Showing posts with label Python. Show all posts
Showing posts with label Python. Show all posts

2014-12-29

Newsletter Updates for December 2014

Chicago, Boulder, NYC, DC, SF, Stanford, London, Stockholm, Madrid, Barcelona, Amsterdam, Dulles, Baltimore, LA. The range of speaking events and business travel over the past quarter almost bewilders, but I’m grateful to get to meet many interesting people and learn about new projects. 

Also feeling grateful to enjoy some quiet time at home with family over the holidays, and I wish very happy holidays to you and yours.

Conference Summaries

Strata NY set a new record with about 450 people attending Spark Camp. There was a spare room, plus an hour break in the fray, so we held an impromptu “Ask Us Anything” about Spark – that has turned into a new kind of open source ritual at Strata confs, especially for handling the more advanced audience questions. Also, Bloomberg kindly hosted a large Spark Committer Night meetup event, their largest to-date.

Manhattan, from NY Water Taxi at Port Imperial
Throughout many conferences and meetup events over the past few months, one demo in particular stood out. David Jonker and Rob Harper from Oculus Info in Toronto gave a talk about Aperture Tiles at Strata NY. Last talk of the show, and quite arguably the best. This open source framework, partly built atop Spark, provides interactive data exploration with continuous zooming on large scale datasets. Highly recommended.

The week after Strata conf in NYC, some of our team found our way slightly south to the University of Maryland, where we got to teach alongside the renowned Jimmy Lin. The week included a Spark Tutorial on campus, plus the initial meeting of the Apache Spark Maryland meetup. Much fun, and we look forward to returning to UMD again soon.

Arriving back to the Bay Area just in time, I caught the launch of the new GalvanizeU program in downtown SF. One challenge that particular evening was getting scheduled to speak head-to-head with the final game in the World Series. That keynote, Data Science in Future Tense, examined some of the near-past and near-future of the field – hopefully indicating some non-intuitive directions. 

GalvanizeU is located next to the Transbay Center, just a few blocks away from the new Databricks office. They provide a hands-on graduate program in Data Science, in an urban setting and working closely with industry partners. Galvanize started in Boulder and is also expanding soon into Seattle. We’re thrilled about our new neighbors.

Home just long enough to take the kiddos trick-or-treating and attend GCPLive, then on to Europe… During a brief visit in the UK, I got to present about the latest in Spark Streaming at the London Spark Meetup: Tiny Batches in the wine (a callback to Don Ho, for those who were born more recently – ideal for getting your luau on). Then on to Stockholm with gracious hosting by Spotify, Ericsson, and SICS.

Good times @ Big Data Spain, Madrid

Madrid came next, for the annual Big Data Spain conf. Noticing a joke painted on the side of a jet at the airport, I had a hunch immediately that Madrid would be lots of fun. I was not disappointed. Our hosts at Paradigma Tecnólogico and Stratio presented an amazing conference, one of my favorites in a long, long time. I was fortunate to give a keynote talk, alongside many other excellent talks, such as from friends at Cloudera and Google BigQuery. I highly recommend Big Data Spain. More about Stratio in a bit…

The beach at El Poblenou, Barcelona
Taking a train from Madrid to Barcelona, admittedly I was missing the former, but Barcelona is a wonderful place. Imagine yourself in Santa Barbara, except that the city is 50 times larger, thousands of years older, and packed full o’ amazing culture. Strata EU was located at a conference center right next to the beach. We held the first official Spark developer certificate exam, plus a large Spark Camp event (25% of the conference attended), a meetup at UPC, and a second iteration of our “Ask Us Anything” about Spark.

Locavore feasting in Catalunya
Business travel Spark-style does not allow much downtime. Effectively one day off during two full weeks in EU. Fortunately that just happened to be during a weekend in Barcelona, the day after Strata concluded. I rented an Airbnb condo near the beach in El Poblenou, then wandered busy Rambla markets, through the crowd surrounding a busker string trio, gathering items to make a small feast. Only in Catalunya.

Amstel River in Amsterdam
A quick stop in Amsterdam, with a very fun talk hosted at eBay with hours of Q&A, then back home. Long enough for a family Thanksgiving feast, then off to DC, Baltimore, and LA. Excellent events and good friends met along the way, particularly the Los Angeles Apache Spark meetup hosted by Rubicon Project. Much appreciated.

Spark

The curiously named Likelihood T. Prior noted on TwitterSpark spark spark spark, spark spark spark spark. #Strataconf synopsis complete. Some went as far as to begin calling “Strata + Hadoop World” by a new name, “Strata + Spark World”. I like the sound of that.

To help keep track of this rocket ride, I’ve begun curating an ongoing list http://goo.gl/2YqJZK of the talks, workshops, etc., related to Spark worldwide. Please let me know if you have events to add.
Speaking of events, recently we began to increase the cadence for Bay Area Spark meetup events. These talks get live-streamed, with the archives published on the Apache Spark channel on YouTube. Databricks also recently announced Spark Packages a community index of packages. The site had to be moved shortly after its launch, due to overwhelming popularity. Good stuff on both the video channel and package repo.

So much news about Spark has happened in the past few months. I’d like to summarize with a few gems collected along the way…
Not least of these items, the Databricks team broke YHOO’s previous world record for the Daytona GraySort contest. That tied for the 100 PB sort on AWS, using 1/10 the number of servers and running 3x faster than YHOO Hadoop clusters. #justsayin

MOOCs

Part of my job involves the curriculum for Spark instruction. Our big news recently is that edX and the University of California will be offering two new MOOCs about Spark, sponsored by Databricks.

The first is Introduction to Big Data with Apache Spark by Prof. Anthony Joseph at UC Berkeley. This comprehensive introduction to Spark, as well as Big Data, is based entirely on Python programming and aimed at developing Data Science skills. This course begins on 2015–02–23.

The second is Scalable Machine Learning by Prof. Ameet Talwalkar at UCLA. This hands-on course focuses on distributed machine learning at scale, based on examples using open data, also in Python. This course begins on 2015–04–14.

Note that some taking Spark MOOCs will have the option to use Databricks Cloud free student accounts. Similarly, we will be integrating use of DBC free accounts into our other Spark training events.

Workplace

Several years ago, I was fortunate to work for a CEO who understood how to leverage a distributed workplace. I studied the management practices involved, and in particular have grown to appreciate ROWE greatly. These practices seem all too rare among early-stage tech start-ups Silicon Valley. However, a few tech firms (DataStax and Typesafe come to mind) have embraced distributed workplace models. Frankly, correlations between effective approaches to gender equality and practices such as ROWE should be on every VC’s radar.

With respect to workplace practices – effective or otherwise – two recent articles caught my attention:
Great words of wisdom about two of the worst anti-patterns for successful tech organizations. The most telling part is the “canary in a coal mine” effect: to watch and see who becomes the most offended by these points. Egregious (sometimes outright hostile) use of email, chat, meetings, etc., and the fallacy of “crunch mode” stand as two of my top determinants for evaluating a company. Right alongside we provide free snacks and meals vs. we offer reasonable health care plans – which somehow turn out to be at odds in far too many start-ups.

BTW, really looking forward to catching Chad speak at GOTO Chicago next May.

Just Enough Math

The Just Enough Math material continues to evolve… Allen and I gave a tutorial at Strata NY, working closely with O’Reilly Media to export content to IPython Notebook within a Docker container for participants to run in the cloud. Rackspace provided the hosting, which in turn was an alpha test for their Nature magazine IPython interactive demo. Welcome to the future of publishing.

Andrew Odewahn and I entered a version of this for the Boston instance of Docker Global Hack Day #2 – frankly, Andrew did like 99.9999% of the work on that one :) Meanwhile, speaking of the future of publishing, JEM provides an example in the new Publishing Workflows for Jupyter by Andrew OdewahnKyle KelleyRune Madsen.

Beyond publishing, we do have some math to suggest… Two papers caught my attention recently:
Oh, and riffing off the “Quantum Algorithms on the Moon” meme from JEM, note that NASA, Google and USRA establish Quantum Computing Research Collaboration such that 20% of computing time will be provided to the university community. In case you have some large data set that’s just screaming to get crunched on a D-Wave. Like you do.

Mesos


Other big news was the Google Cloud Platform Live conference in SF on Nov 1. The message from #GCPLive was largely about containers… in short, the notion of The datacenter IS the computer going mainstream. To paraphrase one comment during the conf: “Customers get locked into host-based patterns, so they struggle with intertwined systems.” Well said. Definitely looking forward to the new GKE service based on Kubernetes.

Other big news was awaiting in London. Namely, the team behind Weave. Recall that the JEM tutorial had been an alpha test for the IPython + Docker + Rackspace + Nature magazine thing? We learned a truism the hard way, with minutes to go before the event started: Docker does little to resolve crucial issues outside of the containers. Enter Weave, handling difficult matters outside the container, such as networking and crypto. Check their blog for tasty insights, e.g., Automated provisioning of multi-cloud weave network with Terraform. Highly recommended.

Speaking of Docker, I really enjoyed this talk by Adrian Cockcroft @DockerCon: State of the Art in Microservices. Especially slides #8–19, product development process.

Speaking of Microservices, here’s a good overview: The Strengths and Weaknesses of Microservices by Abel Avram on InfoQ.

Ag+Data

Continuing on the Ag+Data front, check out the excellent article GeoTrellis Adapts to Climate Change and Spark about how Climate Change analytics drove Spark adoption at Azavea. They integrated Spark and Accumulo to support fast computation of climate impact metrics for DoE, which should be included in the 0.10 release of GeoTrellis.

NYT ran an interactive analysis/visualization, Flooding Risk From Climate Change, Country by Country, which perhaps helps explain Silicon Valley rumors about Google building ferry ports at corporate campuses along SF Bay.

I’m a big fan of Danielle Nierenberg @FoodTank in Chicago. A recent article, How Vegetables Can Save the World, is brief, accessible, and quite to the point. More of that on FoodTank.
Meanwhile, considering the many challenges ahead in Ag worldwide, I’m curious whether some programmable matter could become useful on farms to leverage data? Sort of an asymptote for IoT.

Upcoming Events

Many interesting conferences and other events are planned for the months ahead. Please do check the http://goo.gl/2YqJZK listings. In particular, mark your calendars for:
O'Reilly studio in Sebastopol, for new "Intro Spark" video

Misc.

I’ll leave you with something fun and something epic.

First, the fun – though it’s quite epic in a way: LumiGeekWe make Arduino shields for LEDs, audio-reactive drivers, and custom solutions for architectural and artistic endeavors. Check their installation at the new Galvanize Cafe in SF, and look about carefully for a subtle case of anamorphosis.

Second, the epic – if you haven’t seen it yet, it’s well worth four gorgeous minutes of video: Wanderers by Erik Wernquist, narrated by Carl Sagan. Money quote @1:45: “Herman Melville in Moby Dick spoke for wanderers in all epochs and meridians…”

That's the update for now. See you in Austin, San Jose, and NYC on the event horizon!

2014-07-27

Newsletter Updates for July 2014

Two aspects about leveraging machine learning are largely under-represented in the lit, especially when it comes to production use cases: feature engineering and the comparative evaluation of multiple modeling approaches. To that point, check out “Streamlining feature engineering: Researchers and startups are building tools that enable feature discovery” by Ben Lorica. The article mentions Spark Beyond, which “finds deep patterns in your data.” I was lucky to get a demo of Spark Beyond earlier this year and talk with the principals – and highly recommend taking a good look at their wares. Between the ongoing advances in deep learning and symbolic regression, a direction seems to be emerging … that perhaps one of the more difficult parts of machine learning workflows, namely the feature engineering aspects, could become more automated.

For another great article, check out Including Men in the Conversation About Women by Scarlett Sieber. Among my biggest peeves about Silicon Valley are the “brogrammer” lopsided demographics, and the gender bias which is quite real and nearly epidemic. Our data science teams have generally been quite mixed, why can’t engineering teams in general leave the 19th century behind, let alone stop being so hostile? Not naming names, but two of the SV firms in which I’ve worked in the past five years are both well known and well poised for harassment lawsuits. Taking a stand against that nonsense as an engineering manager is a great way to catch hell, which I’ve gladly engaged before. Another related pet peeve is where one of the same firms was actively pressuring their engineering interns to quit university degree programs. As a behavior for an engineering manager, I find that highly unethical. Some of those who are engaged in these practices know quite well who I’m talking about.

Spark Summit

The big, BIG news last month was … (wait for it) … Spark Summit. All of the speaker videos have been posted – those are probably the single-best resource for learning about Apache Spark. Of course, the big surprise at the conf was the announcement of Databricks Cloud. If you missed the conf, you can watch Ali Ghodsi’s spectacular demo which kicks in at about the 14:40 time marker.

Spark Summit keynote practice, T-15 hours
One surprise learning from the conf was that one product line from SAP generates more annual revenue than all of the other Big Data vendors (HW, Cloudera, etc.) combined. Other pleasant surprises included: Flambo, a Clojure DSL for Spark; and Thunder, for large-scale neural data analysis, which shows some excellent integration of PySpark, SciPy, scikit-learn, etc.

Our training sessions at Spark Summit set some kind of new records. In particular, check out the advanced material for great lectures there. Those who attended the conf received a free ebook preview for the upcoming Learning Spark: Lightning-Fast Big Data Analytics by Holden Karau, Andy Konwinski, Patrick Wendell, Matei Zaharia; O’Reilly Media (2014).

Also, I got to host the Research track of session talks at Spark Summit, which was a real treat. We had a special #geo break-out session following the Geotrellis talk by Rob Emanuele. We will hopefully be expanding that focus in future confs. There were so many other great talks that it’s hard to pick favorites. Even so, I’ll be studying up about two in particular: Quadratic Programing Solver for Non-negative Matrix Factorization with Spark by Debasish Das, Santanu Das; and Distributed Reinforcement Learning for Electricity Market Bidding with Spark by Vijay Srinivas Agneeswaran, Vishnuteja Nanduri. The latter seems almost ideal for integration with recent work on genetic programming.

Stay tuned for the next Spark Summit, which will be held on NYC in early 2015.

OSCON

I’ve just returned from OSCON 2014. What an excellent conference! Check out the content recently posted online: keynotesphotosspeaker slides.

Of course, this event was carefully timed to overlap with the Oregon Brewer’s Festival. Top two picks: Double Latte. by Sierra Nevada Brewing Co.; and Lorenzini Blood Orange Double IPA by Maui Brewing Company. Many thanks to Erin Rasmussen for suggesting about OBF!

Blood Orange IPA
Back at OSCON… one of my favorite Ignite talks was What Science Fiction Can Teach Us About Building Communities by Dawn Foster. Another favorite, speaking of #geo, was a preso/proposal for Open Aerial Map by Kate Chapman.

During the conf, Andy Orem did a video interview where we discussed perspectives and current projects: Ag+Data, Industrial Internet, sketch algorithms, Apache Spark, etc. Andy was the very first editor I worked with at O’Reilly Media, ten years ago. He’s a much better interviewer than I am an interviewee, so I enjoyed learning much through our work together. Also fun to work again with the amazing video team.
"With great power comes some data, plus wrinkled shirts"
The tutorial for Just Enough Math had 50+ people attending, and we got to evaluate an intermediate stage of a new tutorial software platform. For that, I needed to get a bunch of USB drives from Amazon, but the order/delivery #failed. At the last minute our 10 y.o. daughter and I made an emergency run to Fry’s Electronics (she was eager to observe ground zero for nerdliness) … but the only 4Gb flash drives that they had left in stock were Marvel Universe comix characters. Arriving back home, our 9 y.o. daughter was aghast that adults would be receiving comix figures in a lecture :)

The Data Workflows for Machine Learning talk received lots of great responses – as did earlier versions during meetups in Seattle and SF. It become of the “top-shared” slide decks featured on the SlideShare home page. Perhaps that needs to be turned into a mini-book?

new book kiosk
As my last-o’-the-day book signing was winding down, after almost everyone had left the convention center for “nearby locations of beer taps”, a friend mentioned “Hey, look there’s another pile of books – these look different.” So a few lucky latecomers got signed copies of the galley drafts for our new book Just Enough Math, which probably still won’t be released for months – this rev is quite rough :) Oddly enough, the first person to read it looked up and said, “Where are the other O’Reilly books about math?” Indeed.

Sketchy Things

Speaking of Just Enough Math, we’ve put up a companion site for the video+book+tutorial at http://justenoughmath.com/ to provide additional resources and related links:
  • set up a Python programming envon your laptop
  • code+data files for examples in the video+book
  • “gists” that show expected results for the examples
  • links to external resources that get referenced
  • recommended books and videos for further study
  • monthly newsletter sign-up
The tutorial at OSCON previewed a new chapter recently added about sketch algorithms, following from notes at an excellent Foo Camp session led by Avi Bryant. I will be focussing on Spark Streaming use cases for Strata EU in Barcelona this fall, particularly where approximation techniques (think: examples of monoids in action) can leverage both Spark and Cassandra. If you have examples to share of Spark Streaming production use cases in general, I’m eager to build case studies to publish in Radar. Meanwhile, for a great resource about sketch algorithms, check out the archives of the AK Data Science Summit – Streaming and Sketching from last summer.

Card-Carrying Green

A friend recently brought up the topic of navigating questions about extinction and climate change for preschoolers… I’m getting those too; however, in my experience the questions become much better formulated after an additional 5–6 years or so. As a parent, as a human, it kills me to see all the ginormous FUD spewing from the political lobbies for the coal industry, fracking, Monsanto, GM, etc. How about giving ample air time and consideration for some points from the other side?

First off, I’ve mentioned it before but it bears repeating: The Land Institute is a phenomenally excellent resource for understanding some of the insanity and pure tragedy of contemporary agricultural practices, particularly when it comes to monocultures, annuals, hybrids, let alone unnecessary tillage. To paraphrase Wes Jackson, “The plow share has destroyed more options for future generations than the sword.” On a related note, I’ll also point to an excellent article by Michael Pollan, as a forward to Grass, Soil, Hope: A Journey through Carbon Country by Courtney White. Moreover, check out The Solutions Project. That latter site has more substance than perhaps its web-design polish indicates: it’s about the work by Mark Jacobson, et al., on how to power the planet via renewables now while mitigating hurricane damages, etc. One would think that the reinsurance revenues alone would justify a significant investment. In any case, these three links point to the fact that any emerging “dialog of despair” about global warming, etc., is purely FUD. Much can and will be done.

Phylo, the trading card game
I’m particularly grateful to be associated with O’Reilly Media, which provided OSCON attendees with a nice treat in their schwag bags: Phylo, a trading card game. Its gameplay emphasizes endangered species, climate change, food chains, and other environmental pressures. “Phylo is a project that began as a reaction to the following nugget of information: Kids know more about Pokemon creatures than they do about real creatures. We think there’s something wrong with that. Apparently, so do many others.”

In a related development, check out Nerds Without Borders: “We are looking for all sorts of people to help: Engineers, Scientists, Writers, Artists, Dreamers, Activists, Organizers, Fundraisers, Financiers, etc…” Starting with use of IoT sensors and cell phone networks to protect sea turtle hatchlings. Good stuff.

Looking Ahead

Another fun follow-up from Foo Camp and OSCON: getting to talk with Scott Jenson about his work on The Physical Web at Google. Check out his preso, Why Mobile Apps Must Die. The big idea is a kind of “micro-DNS” for low-cost digital tagging of physical items that can be accessed by mobile devices. No app installs required.

In other news, Trafodion was recently released as open source by HP. The name is based on the Welsh word for “transaction”. If you recall about Tandem Computers and NonStop, this product line has a long history of tech innovations – for highly reliable, highly optimized real-time SQL at scale. My uncle retired from Tandem, and lately I’ve spent time with the Trafodion team and am quite impressed. This release brings an interesting new level of Enterprise robustness to real-time transactions+analysis atop Linux+Hadoop. One to watch.

Another to watch closely is The Distributed Developer Stack Field Guide by Andrew Odewahn, Courtney Nash, Mike Loukides, et al. This is a GitHub-based book from O’Reilly. If you see any points in there that need editing, embellishing, etc., then two words: pull request, for the win.

In terms of upcoming events, registration is now open for Data Day Texas 2015, and I’m really looking forward to that. Will be teaching Spark at Scala by the Bay in SF on Aug 8–9, speaking at #MesosCon in Chicago on Aug 21, followed by another Spark course in Chicago on Aug 25.

Flashbacks

I’ll close with a look back to a 1990 Documentary about Cyberpunk. That provides a good summary of what we up to in the early 1990s with Mondo 2000, bOING-bOING, FringeWare, WiReD, The WELL, Turkey City, etc. Tim’s monologue around 15:30-ff is hilarious – both because of his ever-optimistic “There will be mass democracy in the streets” miss, and how much it contrasts with just about every other major point coming true within 25 years. Warning: gratuitous F242 clips, throughout. Time marker 27:11 shows what I was doing as a vendor at many, many raves… Meanwhile, check out a recent bOING-bOING article Alien Autopsy: William Barker on Schwa, two decades later for some of the more astute counterpoint about what was really going on, then and now.


That's the update for now. See you in Chicago with San Diego on the event horizon!

2014-06-26

Newsletter Updates for June 2014

Been quite an interesting month: NYC, SJ/SF, bookended by Hadoop Summit and Spark Summit, with Foo Camp in the midst… much learned, and many excellent introductions.

If you haven’t seen it, this is a gem: Seeing Spaces by Bret Victor, as an evolution of the “Maker Spaces” concept. Another top recommend is A Short History of and Introduction to Deep Learning by John Kaufhold. Money quote: “Learn, don’t engineer feature representations.” Check this review by Mary Galvin at Data Community DC.

For another great source of inspired writings, follow the Matthew Hunt posts on LinkedIn . In this episode of delightfully unexpected connections, Matthew leads us on a path among Pink Floyd, moon cheese, gnome-like cretins, and unlikely heroes for a tale of two Burkes.

   Just Enough Math

The video for Just Enough Math has been on sale for the past month. O’Reilly has a preview video on YouTube, if you’d like to check out a sample. Meanwhile…

I need your help: this Just Enough Math project would greatly benefit from your reviews. Even if you don’t purchase the full video, check the preview and the free sections. We’re eager to hear your feedback, and especially your reviews!

Here’s the thing: on the one hand, if you’re the kind of person who enjoys reading math papers as a fond pastime, this material is probably not for you. There are plenty of other videos in the world, and so many brain teasers, so little time. On the other hand, if you find that math papers tend to be almost entirely devoid of context (which, frankly, many are) and you took math through Algebra 2, and you enjoy seeing some examples, learning some history, etc., then you’ll probably benefit from this video.

There are quite a number of great resources at O’Reilly and other publishers for those who want a deep-dive in any particular area of advanced math applied for Big Data … and the point of the Just Enough Math project is to serve almost like a “hyperlink document” (e.g., old school web pages circa early 1990s) for those other books, videos, websites, etc., along with providing history and case study examples as context.

We’ll be presenting a tutorial based on Just Enough Math at OSCON. Plus, there’s a super-secret discount code for 20% off registration: PACOID

In the Bay Area, we’ve recently launched a Just Enough Math: Machine Learning for Execs and Entrepreneurs meetup. Looking forward to more events through that. Submitted as evidence, check out “How Not to Be Wrong”: What the literary world can learn from math by Laura Miller in Salon.

   Mesos

Some interesting insights about Apache Mesos surfaced in the recent 2014 community survey. And at this point, the list of firms adopting Mesos no longer fits in my browser window. To find out more, check out MesosCon scheduled in August in Chicago. I’m looking forward to talks from John Wilkes and several other experts, and meanwhile will present about Apache Spark running on Mesos. In related news, recently I gave a talk at the Mesos NYC Meetup sponsored by the kind folks at Shutterstock. If you’re in the area check out an intro Mesos talk on 7/17 by Joe Stein at Bloomberg.

✽ ✽ ✽

On a recent camping adventure in Sebastopol, I was grateful to learn about lots of new technologies. One of the more interesting finds was Unbounded Robotics, and I enjoyed a chat with Melonee Wise, CEO. These actually are the droids you’re looking for. Meanwhile, O’Reilly Media is looking for editors, especially in the Data practice area. Got Edit? Join the team!

morning walk in Ceres Community Garden, w/ O'Reilly Media in bkgd

In terms of other interesting technologies… I’ve been hearing memes rumbling about “Big Data is a myth” or “Where are the IoT apps?” Here, that’s where. The part of Nokia that didn’t sell off to Microsoft is handling some of the most interesting fusion of data exhaust that I’ve seen. Case in point, check out Jams, game theory, and equations: the science of traffic for a view of really big data analyzed in real-time. If you’ve attempted to drive anywhere in, say, DC or Austin or Silicon Valley anytime recently during commute times … this is a problem. Money quote: “Then we start to look at the car’s sensors. We start to know the weather before the weather authorities do, because we can see which cars have their windscreen wipers or their headlights switched on.” Orders of magnitude larger than your favorite social network or ad exchange.

Meanwhile, my favorite IoT app so far is clearly this: sharks tweet as they approach the shore of Western Australia. Would be great to see more technology applications like that!

   Minecraft camp

Speaking of Foo and other camps, I’ve got two kiddos currently in iD Tech Camps --learning Minecraft and Scratch, respectively. These courses tour around the US and are highly recommended. We could learn much from their teaching approach, to benefit professional workshops for adults as well.

To follow-up on the Minecraft + Quantum theme from previous posts, here’s a good video of Seth Lloyd explaining Quantum Machine Learning. Why does this seem to call back to the Real Genius movie?

   Ag + Data

O’Reilly Strata recently carried a story about how Farm data could be worth billions, related to the Ag + Data post on O’Reilly Radar. Much is happening in Ag data and other consumers of remote sensing products – particularly with respect to recent changes in satellite regulations. However, my favorite recent Ag story is about the Purdue Improved Cowpea Storage (PICS) bag. Brilliant work.

Overall, much of the interesting Ag+Data tech seems to be coming from (or through) Chile… and a new phrase has emerged: Chilecon Valley.

   Friends in the News

Congratulations go out to Robby Garner, competing with the JFRED Chat Server in Turing2014: 60th anniversary year of Alan Turing’s untimely death. Many years ago, Robby and I worked on a primordial version of JFRED. That played “customer service agent” for the FringeWare online bookstore. Circa 1998 we ran the bots on BBC “Tomorrows World” for a live televised Turing Test, which is some of the  most fun I've ever had in network engineering. More recently, Hubot-based chatbots are being deployed for devops and other engineering teams, such as the Shep chatbot used by engineering at O’Reilly.

Also, check out the new Big Data Analytics Beyond Hadoop: Real-Time Applications with Storm, Spark, and More Hadoop Alternatives by Vijay Srinivas Agneeswaran. This is a deep-dive into design patterns and frameworks for large-scale analytics beyond Hadoop.

Got to meet lots of people interested in using Spark at the recent Hadoop Summit in San Jose. One of the Community Choice Awards at the conference went to “Demo: Building a Unified Data Pipeline in Apache Spark” by Aaron Davidson from Databricks. Eager to see the slides published for that. Also at Hadoop Summit, Xiangrui Meng gave an excellent talk about the MLlib – the tech roadmap and integrations, and especially emphasizing about how to leverage sparsity in your data.

Meanwhile, friends at Zementis have recently released PMML support for Python, with a project called Py2PMML. In particular, there’s integration for scikit-learn. I wonder how long before PySpark + MLlib joins that list?

   Joaquim on the Moon

As many of you know, given enough beers I become fond of talking about dropping large complex arrays of sophisticated equipment into the polar dark craters on the Moon. In recent convo over drinks with people who calculate the costs of such an operation, for a living we surfaced a interesting price tag for that kind of venture: approximately $15B. In terms of how much the US spends on the Department of Defense, that’s about 8 days’ worth. Think about it. Who wants a term sheet? Meanwhile, the subject got me thinking of Kubrick films, particularly the 2001: Space Odyssey set production, an engineering feat in itself.

That's the update for now. See you in PDX and ATX, with Chicago and San Diego on the event horizon!

2014-05-26

Newsletter Updates for May 2014

Been quite an interesting past month or so: DC, Austin, SF, Ann Arbor, Atlanta, Seattle… with hopefully much learned from those travels, plus many excellent events and introductions.

Meanwhile, I learned much from this gem, Therbligs for data science: A nuts and bolts framework for accelerating data work, by Abe Gong. Looking forward to seeing more about Therbligs from Abe. Definitely tune in to Welcome to Intelligence Matters, a new series by O’Reilly exploring current issues in AI, with Beau Cronin as lead correspondent. Another recommended gem is Genomics Crash Course for Data Engineers by Allen Day – that's at the intersection of Genomics and Big Data, for which I have seen an uptick recently.

Just Enough Math

Allen and I have been working to complete our new O’Reilly book, Just Enough Math. The video is in post-production now, and the book is half through second drafts – we are closing in! Some of that material will be previewed in the upcoming workshop Machine Learning for Managers:
O’Reilly will host a free one-hour webcast, Computational Thinking, Just Enough Math on Wed, Jun 4, 10:00am–11:00am (Pacific). Please join me there. The webcast will help publicize a tutorial based on Just Enough Math at OSCON in Portland on Sun, 20 Jul, 9:00am-noon. As a special offer, use the code PACOID to get a 20% discount on OSCON registration. Our tutorial will preview a very new thing at O’Reilly: converting book+video content into interactive tutorials using Docker + IPython Notebook + Vagrant + Git for a cloud-based next-generation content platform.

Speaking of Docker, one of the more interesting start-ups that I have run across recently is Resin, using Docker and Git to containerize+push apps on IoT devices running embedded Linux. Brilliant work.

UCB Initiation Ritual: cousins circa 1968, near Atascadero

In other news, I am thrilled to announce a partnership with Databricks, where I’ve been working to help develop an instructional program that introduces Apache Spark. As you can see in the photo above, the ceremonial ritual for teaming up with UC Berkeley is a bit arduous, but well worth it. Yes, you heard correctly … a Stanford alum saying “Go Bears!”

Our first course in the series is Databricks Hands-on Intro to Apache Spark, an introduction for developers working in Python, Java, and Scala. We have several of these workshops scheduled:
Spark is approaching the 1.0 release at Apache, with new support for SQL. Overall, one of the best presentations that I’ve seen recently about it was Spark at Twitter by Sriram Krishnan, Engineering Manager for Data Platform at Twitter.

The agenda was posted recently for Spark Summit 2014, in SF on 30 Jun - 1 Jul. As another special offer, use the code Paco2014 to get a 15% discount on Spark Summit registration. Highly recommended, and I hope to see you there.

Mesos Updates

Speaking of BDAS and the Berkeley Stack… there have been lots of developments in the Apache Mesos world. One of the best talks ever about Mesos was Improving Resource Efficiency with Apache Mesos by Christina Delimitrou, a case study about Quasar usage at Twitter. Also check out Mesos Elastically Scalable Operations, Simplified by Niklas Nielsen and Adam Bordelon, presented recently at ApacheCon 2014.

The other big news is that #MesosCon, the first Mesos conference, will be held in Chicago on Aug 21. Definitely see you there! Companies interested in sponsoring the conference – please inquire.
I’ve create a new workshop called Cluster Compute App Integrations about building end-to-end apps for Big Data. The workshop leverages Mesos based on the https://elastic.mesosphere.io/ service in the cloud, along with Spark, KNIME, etc. Hint: this involves teams competing, and it is turning out to be quite a popular course. We have upcoming dates lined up:

Agriculture + Data

Did you know that agriculture provides a livelihood for 40% of the world’s population? Or that agriculture consumes 70% of the world’s freshwater in aggregate? That figure is expected to reach 89% by 2050. Or have you heard that Havana grows 75% of its own food based on urban agriculture?
Last month I wrote an O’Reilly Strata article, Ag+Data, about those topics and more. The article introduces a whitepaper, Agriculture + Data: Outlook 2Q14, that we recently at The Data Guild to explore these issues in greater depth. Many thanks to Bill Worzel, Brad Martin, and others who helped on that!

Evolutionary Algorithms

Recently I gave a keynote talk at the Genetic Programming in Theory and Practice conference, which hosted each year at U Michigan by The Center for the Study of Complex Systems. They are the experts in GP; I was merely there to add a few perspectives about machine learning and Big Data. What a wonderful conference. Got to speak at length with Lee Spector at UMass Amherst and Hampshire College. Lee and his grad students have been working with a Clojure-based language called Push, in which evolutionary programs are expressed.

What kinds of optimization problems respond to evolutionary pressure? Definitely not the kinds that one typically finds solved by machine learning. That is where GP approaches come in. In general, there was a lot of discussion about symbolic regression as a general rubric, also some exceptionally interesting work on use of Pareto optimal fronts for model archives (which I’ll be added to my ML bag o’ tricks). In particular, great work from Theresa Kotanchek and Mark Kotanchek at Evolved Analytics. Their software effectively leverages Pareto optimality to select exemplars when models diverge, which I find to be a fascinating alternative to what other disciplines might attempt to resolve through sample. Brilliant work.

Also got to talk with Bill Tozier, author of Answer Factories: The Engineering of Useful Surprises, and viewed some astounding work in HeuristicLab, an interactive framework from HEAL. Think: evolutionary IDE. Another excellent tip was to check out Modeling global temperature changes with genetic programming by Karolina Stanislawska, Krzysztof Krawiec, Zbigniew Kundzewicz.

✽ ✽ ✽

Didn’t get to mention yet about Atlanta, but I really appreciated meeting many wonderful folks there. You’ll be hearing more about upcoming Atlanta plans soon! Also, there are workshops and meetup talks planned now for: NYC, SV/SF, Austin, Chicago. Next up after my current week in Seattle comes Hadoop Summit, on 3–5 Jun in San Jose. Hope to see you there!

-alaVoid Distribution

Misc. Inspiration

In closing… Those who have known me for, well, for the past 20-odd years or so will be familiar with the following: a 21st century artist named William Barker, formerly acclaimed of Schwa Corporation has a new endeavor called -AlaVoid Distribution. Definitely check out his new shop on Etsy.

2014-04-02

Newsletter Updates for April 2014

If you have not seen Data Science Folk Knowledge by Krishna Sankar, that is packed full o’ gems about Machine Learning.

Following up on more follow-ups from Strata SC 2014, I’d like to point to an excellent article: 5 Steps to Thinking Like a Designer in Machine Learning by Kevin Dalias:
We’ve all heard the saying that a data scientist is a cross between a statistician, domain expert, and machine learning hacker, but in today’s landscape, that falls short. A good data scientist needs to be all of the above and also a great designer.

Great points made there. Thanks to a teaching fellowship as a grad student many years ago, I got to stick around a couple extra terms and take a Design Communications program. That program evolved substantially, and of course there’s now Stanford d-school too. I'm grateful for those experiences and (taking a cue from Kevin Dalias) perhaps some formal exposure to design helped shape my later career path. Most certainly we emphasize design thinking at The Data Guild among our core values.

Speaking of Stanford, a recent report entitled Researcher reveals how “Computer Geeks” replaced “Computer Girls” pegged a ginormous criticism I have had about Silicon Valley:
This stereotype of the lone male computer whiz is self-perpetuating, and it keeps the computer field overwhelming male. Not only do hiring managers tend to favor male applicants, but women are less likely to pursue careers a field where feel they won’t fit in … as late as the 1960s many people perceived computer programming as a natural career choice for savvy young women.

As a parent of two girls who have dived into Minecraft eagerly along with their friends, I’m hopeful for a more balanced future. Perhaps some encouraging messaging is a good step toward that. Wired recently reported In a First, Women Outnumber Men in Berkeley Computer Science Course. Some refer to the Smithsonian report as counterpoint. Nonetheless, the male bias in tech is unquestionable... and that lone wolf thing was always a really terrible idea.

O'Reilly Video Studio, Sebastopol – meerkat crossing

Just Enough Math

Got to spend much of last week in Sebastopol working with the remarkably talented folks at O’Reilly VideoAllen Day and I have been busy writing our new book, Just Enough Math – and are now developing a video and much more to go along with that. Elevator pitch: advanced math for business people, to understand how to leverage OSS frameworks for Big Data. We pick up from a prereqs of: High School Algebra 2some Python, plus the experience in business to recognize why you need to leverage data. Let's explore how.

For example, you've probably heard much about graph query engines and perhaps read about graph use cases … how much graph theory did you get exposed to in school? Given that so many people stop at calculus, graph theory is perhaps rare among b-school topics. Would you feel comfortable working from a use case – in the sense of an HBR-styled business framework – leveraging large-scale graphs to build a high-ROI app? We think the answer is “Yes.”

Each morsel of advanced math gets introduced through a clear business use case, some historical context, lots of illustrations, and small snippets of Python code that you can cut&paste. In addition to integrating text + video + code, we are leveraging an instructional rubric called Computational Thinking. Stay tuned! Meanwhile, we’re scheduled for a Just Enough Math tutorial at OSCON in Portland the week of July 20th.

Compelling Projects

While I’m out teaching workshops around the country, I enjoy many opportunities to meet amazing people and hear about their projects. I’d like to highlight in particular about Mike West in Austin, and his recent post People Analytics Junto (public community):
We are a loosely connected group pressing forward, for the benefit of humanity, on the following topics: 1.) Exploring innovative ways data can be used to solve people related problems or make better people related decisions in organizations. 2.) Seeking understanding of organizations, and people in organizations, through Behavioral Science - Sociology, Labor Economics, Social Psychology, Psychology & Operations Research… 3.) Applying Big Data, Machine Learning, Artificial Intelligence, and related disciplines to Human Resources.

I am fascinated by use of Big Data and machine learning for HR. Having worked closely with HR in several organizations, having hired lots of people into Data Science and Engineering roles… I struggle to point to instances when we really leveraged data much – other than calculating salary targets for new hires or attrition rates.

The point is not to automate HR. Rather, the point is that most organizations spend most of the revenue on people (which makes sense) so why not invest in data insights there?

Do You Need or Want to become a Data Scientist?

KDnuggets recently ran Part 3 of 3 in my email Q&A interview ... aimed at candid career advice to people wanting to move into Data Scientist roles. Arguably a bit over the top, and in reaction to being exasperated by "Read this and begin calling yourself a Data Scientist" puff pieces. Many thanks to Anmol, Gregory, et al., @KDnuggets. This came after part 1 and part 2 about Apache Mesos following Strata. I'd like to publish some snippets here – not about what I said, but about what people began discussing.

A comment by Data Science London:
++1 “Product Mgmt. in SV is almost antithetically opposed to effective use of data” … Data Products nuke mgmt layers
Another comment by Data Science Retreat:
“Actual work in Data Science entails having to speak truth to power (not fun, but the essence of the role)”
Several criticisms were much appreciated, and hopefully that helped spur more dialog... Daniel Tunkelang:
@paix120 @BecomingDataSci I agree that @pacoid is overcompensating a bit to counter the proliferation of be-a-data-scientist-quick programs.
Data Science Renee:
.@dtunkelang @paix120 @pacoid hm could be. Just seems to take an unnecessarily discouraging tone.
Followed soon after by great perspectives which are recommended reading.

Plus some wise words stated much more succinctly that I could... Gregory Primosch:
a #datascientist is not a magical unicorn http://goo.gl/ytHzT4  
Andrew Musselman:
@GPrimosch @pacoid I started saying instead of hiring unicorns you should hire horses and narwhals
All of that discourse was illuminating to see. Well said, much better than I did. And, as Charlie Greenbacker pointed out:
It’s also a great rebuttal to all the articles claiming “data science” will soon be automated. #WishfulThinking
Indeed. Not to be overly flippant, but my hunch is that Data Scientist roles will become fully automated at about the same time as HR professionals and BoD meetings.

Meanwhile, IMHO some of the wisest words on this subject come from Nick Kolegraff, Dir Data Science @Rackspace: Do you need a data scientist?

Open Source Updates

There’s an excellent Mesos community update summary on the Apache Mesos site, along with Cassandra on Mesos integration, and a new task load simulator framework for cluster performance analysis. I also recommend an excellent preso from Claudiu Barbura @Atigeo about tying together Mesos, Spark, Cassandra, etc. There are building blocks for datacenter computing.

I’ve been working with Apache Spark lots more lately, particularly diving into PySpark. Check out the recommended Spark SQL: Manipulating Structured Data Using Spark … those are Spark workflows, not Hive fronted by Spark. Very nice work. Also, make plans for Spark Summit 2014.

Other recommended conferences with recent announcements:
Backyard in bloom – site of a new Google campus

Upcoming Events

Lots of plans to be out on the road during April/May this year. I hope to get to talk with you there! Here’s a summary of upcoming meetups and workshops, including new material. We will have drinkups plus office hours in most of these cities – probably adding more meetup talks too:
Misc. Inspiration

Speaking of Big Data apps, here’s a good one: simulations show that in the context of hurricanes Sandy, Isaac, and Katrina, wind farms disrupt the outer rotation winds so much that the storms do not even have enough energy to destroy the turbines – let alone their damage after landfall. That would effectively reduce category 5 storms to category 2 on a Saffir–Simpson scale. Similarly, storm surge decreased substantially in simulations – up to 79% for Katrina.

Consider that estimates for protective seawalls run in the $10-40B range, per city installation… seems like reinsurance companies could start underwriting turbine farms to cut their costs massively, not to mention generating electricity. I've heard Prof. Mark Jacobson present, and his work in general is highly recommended.

I'll leave you with one of the more interesting kinds of archival data that I’ve seen in a long while… Years by Bartholomäus Traubeck, a record player that plays tree rings.

That's the update for now. See you in DC, Austin, Ann Arbor, and Atlanta, with Seattle and NYC on the event horizon!