The split-apply-combine paradigm in R

Last night at the DC R Users meetup, which was our largest meetup to date, I gave an introductory presentation on data munging, and spent a bit of time on the split-apply-combine paradigm that I use almost daily in my work. I talked mainly about the packages plyr and doBy, which I use a lot now. David Smith posted a link on the Revolution blog to this article by Steve Miller, talking about the virtues of the data.table package for doing “by-group processing”. It got me thinking about changing my workflow yet again and engaging this package in my computational workflow. I also noticed that Hadley Wickham tweeted that he wants to make plyr faster as well in the near future, which will of course be a very welcome development.

3 comments

Hi Abhijit, is the presentation available for distribution?> Thanks.

Abhijit says:

March 2, 2011 at 10:40 AM

Yes, it is at http://files.meetup.com/1503964/DC-RUG-Meetup-Feb-24-2011.zip

Reply

As a person who is just venturing into R from SAS …… this is REALLY GOOD and helpful.

Thanks so much for sharing !!!!

Stat Bandit

Musings on statistics, computation and data research

The split-apply-combine paradigm in R

3 comments

Leave a comment Cancel reply

Share this:

Related

3 comments

Leave a comment Cancel reply