Showing posts with label Spark. Show all posts
Showing posts with label Spark. Show all posts

2015-05-31

Newsletter Updates for May 2015

If you haven’t already been listening in to the O’Reilly Data Show Podcast hosted by Ben Lorica, then by all means do not walk, run to check it out! Episode linked above is with Anima Anandkumar @UC Irvine, recommended previously here, re: tensor analysis. Similarly, recent collaboration among David Gleich @Purdue, Austin Benson @Stanford, and Lek-Heng Lim @ UChicago, using tensor analysis to resolve hard problems in higher dimensional Markov chains, resulted in “Spacey Random Walks”. Punchline on slide #17. I detect a trend…

Confs

Three months and so much travel since my previous post: to paraphrase Ricardo Alberto Fernando Ricardo y de Acha, I’ve got some ’splaining to do :)

Highlights include:


Really enjoyed the track chair gig for Data Science. My two favorite talks, both highly recommended:
That was a busy conf, indeed! You can tell since no time was carved off for sacred pilgrimages to Wassail NYC cider bar, Mast Brothers cold brew chocolate, or Blue Hill Farm. Instead my flight left the following day for…


My first time to Brazil. I’m hugely impressed by the developer community there in SP. Most inspirational quote from the conf: “A team is like a symphony, not a factory.” by Randy Shoup.
Great sessions on Spark, Docker, and much much more. Fortunately, we did have time to sample the local cuisine… e.g., Italian food made with tropical ingredients, or my favorite meal of the week, Pirarucu steamed in banana leaves, with a local favorite recipe for pumpkin soup. Then back to the US, before one could even say “alfajores” – assisted on a course at Stanford, then back to NYC…

Pirarucu – Brasil a gosto, São Paulo


Organized by Jeremy Freeman and crew. Truly excellent: top neuroscience researchers in the world gather for a hackathon (i.e., coding together). Perhaps a few savvy Finance people lean in too, eagerly drooling over results that apply for their high-dimensional, time-series, non-linear correlations work as well… in what VCs insist on calling an “ecosystem”. Or something. See these excellent notes. My favorites:
  • Michael Dewar : streamtools real-time analytics from NY Times
  • Olga Botvinnik : flotilla - Py package for iterative machine learning analyses
  • Eiman Azim on motor circuit function … with actual electronic circuits to emulate what Columbia discovered through ablation studies about cerebellum connections
  • Brendan Lake teaching computers to scribble characters like humans, if you want some really interesting use cases for Deep Learning
We worked together the following day, between tutorials, to build an online platform for submitting algorithms to run against standard neuroscience data sets. This hackathon literally was research. If you want to understand more about Big Data being used expertly in life sciences, attend CodeNeuro!

Lower East Side, Manhattan – from New Museum roof


Many thanks to Marilyn Waldman, Claudia Imhoff, and crew for a fantastic Spark tutorial at CU Leeds Business School! Followed by excellent convo via webcast with several hundred of the top BI analysts in the world. Then on to Boston for…


Matei and I were multiplexing between these two conferences so much, UberX in oscillation overthruster mode, that we didn’t even see each other. Even so, lots of great Spark talks in Boston that week! Meanwhile, kind hotel staff redirected me toward The Barking Crab for dinner, and a good friend introduced L.A. Burdick. Also, there was lots of excellent hard cider in the area. Along with all that Big Data conf talk stuff. Then a red-eye flight took off for Europe…

Boston Seaport, by water taxi


Many thanks to all the work by Amparo Alonso-Betanzos, David Martínez-Rego, and colleagues for organizing the Spark tutorial at A Coruña. I’d never visited Galicia before, a place with lots of rain and people with red hair playing bagpipes amidst rolling green hills (no en UK – pero en España) … a place where there are software companies next to world championship surfing competitions and albariño vineyards (no en Santa Cruz, California – pero en España) … a place where they speak a language close to Portuguese – not so unfamiliar, right after São Paulo! Excellent other talks, along with a Spark tutorial by Juantomás García. What incredible people, Computer Science excellence at the university, and oh such good sea food. My keynote was broadcast on Spanish television – that’s a first! Use of runways in this corner of Spain is quite abbreviated, so we catapulted next for…

Frente a la Torre de Hercules


See a good summary of the conf online. Spark Camp this time had 12% of the conference attending, whereas before we’d been trending steady at just over 8%. Many thanks to all who participated. Great to meet so many people enthusiastic about using Apache Spark! Also got to host the Hadoop and Beyond track. Which, oddly enough, was mostly about Spark. My two favorite talks:
Plenty of hard ciders sampled while in London. Whenever visiting near West London, I make a point to drop by Princess Victoria. Also, got to see some friends at UCL and the Barclays incubator program. Then back to the US via customs in my former part-time home city, Vancouver…

Chicagoland


Dean Wampler, et al., have been hosting great Spark events in the Windy City. Always a treat to visit. We had a good turnout for the Spark tutorial at GOTO, one of my favorite software conferences. Walking alongside the Tribune Tower after my tutorial, I noticed that its walls contain rocks from other famous buildings all over the world. With labels, like an inverted museum. Check it out when you visit. There are also rumors of a cider bar being built, ahem, soon. Then a flight back to the Bay Area…


Serendipity. Got invited by a dear friend, Donna Kidwell @Webstudent to join this conf about the future of education, a collaboration among Future Learning Lab, H-STAR, EdCast, etc. Many thanks to Oddgeir Tveiten and crew.

Most discussions focused on experiences with MOOCs – from highly successful examples, e.g., Intro to Robotics by Peter Corke @QUT, or Trust Academy @Salesforce by Masha Sedova. However, the overall themes transcended MOOCs, asking the question of what comes “After Gutenberg”, how peer evaluation is transforming education at scale, and … Peter Norvig’s point that when so much learning material is available (e.g., via Google) the problem becomes a matter of how do you get people to want to interact with it? Generally, the social context of learning becomes key.

Along similar lines, Michael Shanks stressed that – in contrast to traditional academe that tends to decontextualize learning – more contemporary advances are focusing on how to contextualize, locally. That’s part of the essence of interdisciplinary work, e.g., Data Science. Also, delighted to meet Keith Devlin (our new neighbor) with amazing work in math education using Minecraft, etc. (sound familiar?) And was very fortunate to meet teaching superhero David Conover, who uses game design to teach topics like IoT in an at-risk high school in Austin. Brilliant.

Of course, we’ll be inviting the whole lot to propose talks for Strata! Speaking of advances in learning platforms, check out the recent beta site and related article Embracing Jupyter Notebooks at O’Reilly by Andrew Odewahn. Par example, Data visualization with Seaborn … that is integrating use of IPython/Jupyter, Docker, Thebe, etc. Brilliant++.

Data Science

Another recommend: a new series of ML-related interviews by David Beyer, beginning with my friend and colleague Reza Zadeh @Stanford – on the evolution of ML, deep learning, Stanford ICME, and Apache Spark.

If you haven’t seen the news, Nature banned use of p-values … #finally  Note that Fisher did not intend for p-values to be (ab)used that way. So I consider these tests to be truly excellent, for identifying intellectual limits. Related: The Nine Circles of Scientific Hell.

On the subject of pseudoscience – appalling to see recent and ongoing unscientific gaffes by people who should know better, e.g., Neil deGrasse Tyson about GMOs. In great contrast, antidotally, I’d point toward this gem – and please read it at least twice: Die, selfish gene, die by David Dobbs. So glad to see Dawkins getting served … #finally

Neil, Richard: just because you’re published, doesn’t mean you’ve become superior thinkers. Please keep your day jobs, respectively #thankyouverymuch

Speaking of ’splaining to do… another gem is Visualization Explanations @Setosa. Extra points if you grok the callback in their name.

If some of these items are what you’ve been talking about recently, you may just be a Data Science instructor … or interested in becoming one? Check out Become a Data Science instructor @ Galvanize (Seattle, SF, Boulder)

Spark

También se recomienda: una excelente introducción en español a la Scala y Apache Spark por Isra Gaytan. The Latin World has been busting a move lately on Spark, #justsayin

A Coruña – from Playa de Oza

Adatao published an excellent article about how they anticipated the inflection point for Spark adoption. Check out the cost curves. Brilliant.

Big news from AMPLab: Keystone.ML is released as open source, to make the process of constructing complicated machine learning pipelines easier. Good stuff from the eponymous Evan Sparks, et al.

I’ve got some graph analytics talks coming up… and was excited to see a streaming/incremental SSSP impl for Spark.

Also, speaking of the Apache Spark Developer Certificate, we’ve got another new neighbor: ORM + DataStax partner on C* cert. Collect all three!

Meanwhile, the big BIG news is Spark Summit 2015 coming up next month in SF. Use the discount code SparkSummitPC25 for 25% off registration. Not retroactive, but nice try :) Followed by Spark Summit EU in Amsterdam, this autumn. Spark it up!

IoT

Solid Conference is coming up again soon! Highly recommended. As an appetizer, check out this excellent article by Cameron Turner @The Data Guild: Caltrain Quantified - An Exploration in IoT which we could hear in our previous backyard every morning starting at about zero-dark-thirty. Now that former backyard has become the new GoogleX building, and SciFi tech experiments compete with the trains for attention.

For IoT in practice, I’m totally stoked to see: Surfers on acid… What an excellent application. And, culturally not far off that mark, here’s an interesting take on marine plastic: Net+Positiva.

Ag + Data

I’ve really been enjoying Biocoder News in quarterly installments, some of my favorite new articles in the world. Period.

On that note, I’m thoroughly ecstatic to announce that I’m moving to O’Reilly Media full-time. Even so, I'll stay involved with Spark and Databricks, assisting on Spark Summit, etc. We’re moving the family to a tiny farm, an old apple orchard that really needs some tending. Perfect as a research station for Ag+Data.

The Tiny Farm – redwoods 30m tall, planted 65 yrs ago by previous owner

In highly related news, check out How to Grow a Forest Really, Really Fast, about fantastic work by Shubhendu Sharma. I’m eager to try this out.

One of the top intellects of the early 21st century, Paul Stamets, had some excellent coverage: He Holds The Patent That Could DESTROY Monsanto And Change The World!  See also: BioMason  and Ecovative Fungi, FTW – and mycorrhiza in particular, as Mohamed Hijri explains quite succinctly.

Ag-related tech approaches in SV have become largely derailed by asinine priorities dictated by Monsanto – more about taking over hedge funds on commodity trading globally, than about feeding anyone. Perhaps the best analysis that I’ve read recently – and certainly one of the best books that I’ve ready recently – is the highly recommended The Third Plate by Dan Barber. I learned about that via Gastropod – where Cynthia Graber and Nicola Twilley consider food through the lens of science and history. Brilliant.

There’s been a terrible drought / water crisis in Brazil – largely exacerbated by transnational corporate interests. This was weighing on my mind as our flight landed in São Paulo. I got to speak with friends there who are working on Ag+Data analysis, very good to see.

As predicted: Finance is driving California water into the dust… take a moment to consider the jump in almond production versus the temporally co-located jump in variance for snow pack levels. That’s the tip of the iceberg for the near-term shape of major political battles brewing in California. To wit, some of our local mafia have become known under the more apt monicker of Oligarch Valley. While Fox News, et al., promotes the Israeli approach of desalinization at scale, many people who can actually think for themselves question the impact of that approach, and recognize what an utter environmental disaster it could produce. This is not an area of judgment where one gets to call #oops as an excuse, regardless of which side the local mafia may be taking.

Industry Insights

“Software eats the world” is a catchphrase used by A16z. While I slightly agree with the title from this Datanami article, How Machine Learning Is Eating the Software World, its conclusions are pretty much the opposite of what we’ve observed with Apache Spark use cases in the field. Don't get me wrong – Reynold is a good friend, and IMO one of the most talented people working in distributed systems today. However, I have a hunch that the reporter munged the line.

Two key reasons why organizations adopt cloud-based notebooks are (1) to reduce their need for DevOps people to run clusters; and (2) to reduce the need for programmers to assist business people with queries for insights Big Data. Done and done. In other words, domain experts trump all in Data Science applications, while application developers (in relatively large supply, but relatively expensive) and expert systems engineers (in relatively short supply, extremely expensive) both become less of an existential bottleneck for new ventures. I’ll let you do the math on that one.

Some of the themes that I’ve been researching and illustrating over recent years include: Functional Programming for Big Data, Approximation Algorithms, Tensor Factorization, etc. Recognize that each of these point toward less emphasis on developers leveraging APIs, and meanwhile more emphasis on domain experts leveraging simple-to-use frameworks. That’s the bottom line of Apache Spark. Meanwhile, I have no doubt that A16z will continue to rake in loads of money – some of their partners are well-connected billionaires – just perhaps not as a consequence of their thesis. That ship is already sailing. Off, perhaps, toward the oh-not-so Great Pacific Garbage Patch.

Diversity

Speaking of VCs in SV… TechCrunch analysis recently found that female founders nearly doubled in 5 years. Par example, check out the recent Women in Data: Their Work and Achievements.

Meanwhile, I thoroughly enjoyed this gem by Karin Rubin: How women are conquering SP500… My feelings about the overall ethics of algorithmic trading are arguably mixed. However, if it’s going to happen, why not guide it based on diversity, since that demonstrates a #winning strategy?

Fun Stuff, friends in the news…

Check out Lumo Interactive Projector by Meghan Athavale and crew. It’s an interactive floor projector, transforming a floor into games that kiddos can design themselves.

Also, this bit Our Coming Robot Overlords about David, Amanda, and Zeno Hanson – friends back in Texas.

And, what William Barker called one of his most honest interviews, ever.

Upcoming Events

Will just leave you with…


This article. Wonderful, on so many levels.

2015-02-28

Newsletter Updates for February 2015

Not so much travel recently – Austin was my only trip this quarter so far. We’ve been heads-down reworking instructional materials to highlight what you can do with cloud-based notebooks. To learn more about that, check out the new Databricks newsletter.

Snow near Cold Springs, California

Meanwhile, my family gets to enjoy some time this weekend in a cabin near Yosemite, during an increasingly rare event here: lots of snow! Recommend: we always try to drop by our favorite mile-high restaurant, Mia’s, for excellent Italian cooking in the mountains and even homemade limoncello.

Strata

Of course, one of the other big reasons for keeping close to home lately was our biggest event of the year, Strata + Hadoop World in San Jose. Here’s a link for the published speaker slides and videos, along with an excellent summary of the Hardcore Data Science day by Ben Lorica.

About 325 people attended our Spark Camp tutorial. Oddly enough, that’s the same ratio of total conference attendees that we had at Spark Camp in NYC last fall. I also got to host the new Spark in Action track. One eye-opener in our track was the Tencent talk, where LianHui Wang presented about their experiences running an 8000 node Spark cluster in production. So much for FUD claims that Spark doesn’t scale ;) When asked how Tencent can build substantially larger clusters than what YHOO has reported, LianHui replied wryly, “They do not speak Chinese.”

StackOverflow analysis of Spark by Donnie Berkholz @RedMonk

One of the other Strata talks that I really wanted to catch: Tensor Methods for Large-scale Unsupervised Learning: Applications to Topic and Community Modeling by Animashree Anandkumar @UC Irvine. For more details, check out her video. 

In particular note the experimental results at the 42:46 mark, along with slides for a related talk. There is even more background in the recent papers: Guaranteed Non-Orthogonal Tensor Decomposition via Alternating Rank–1 Updates; and Tensor decompositions for learning latent variable models.

The gist of this effort is about using graph moments, assuming priors which then help make tensor decomposition tractable. This material will flex your advanced math agility as it flies through linear algebra, graph theory, statistics, and optimization for some startling implications. While the immediate research is about latent variables for community detection (think: Facebook) these techniques have implications on a much broader range of industry optimization problems. Note that the outcomes are in contrast to work by Jure Leskovec, et al., @Stanford. Another excellent Spark-related talk at Strata that referenced work with tensors was Hadoop as a Platform for Genomics by Allen Day @MapR .

Looking Ahead

Why tensors? Recall from 18 months ago, “I give Hadoop three years before it gets displaced.” At the time that prediction drew some flack. Now that we’re halfway to the predicted time, note that during the past three Strata + Hadoop World conferences there have been numerous remarks to rename it Strata + Spark World. However, the general insight drives a bit deeper…

My question here is, “What is the business case for developing custom apps atop a Hadoop platform?” When I examine industry use cases for Big Data frameworks, there are a few general categories:
  1. ETL
  2. data warehouse replacements
  3. data exploration and reporting
  4. analytics in depth, leading toward streaming
The first category is relatively well-understood, leading toward general purpose solutions. On the start-up side of the spectrum there are great solutions emerging such as ETLeap, Alation, and arguably examples such as Epic in medical data exchange. On the established side of Enterprise IT, incumbents such as Informatica have been aggressively partnering and expanding the scope of their integration. That begs the question of whether firms would continue to build rather than buy?

The second and third categories are the devil-you-know, as continuations of DW and BI respectively. SiliconAngle had a good article recently along these lines, The cheat sheet to following Big Data’s money trail by Suzanne Kattau.

My hunch is that in terms of the second category, Cloudera, Hortonworks, etc., will be forced to pivot toward vertical applications sooner than later to sustain their growth, and will likely buy up smaller analytics vendors along the way. That puts them on a collision course with incumbents Oracle, IBM, Teradata, SAS, etc., where both ends of the spectrum race toward resembling each other. In other words, the DW king is dead, long live the DW king. Expect either some contractions or M&A activity as a result. Not much news there.

The third category, effectively a BI displacement, gets a bit more interesting. I gave a keynote talk at Data Day Texas in Austin in January, A New Year in Data Science: ML Unpaused. The gist is that two aspects of the BI displacement – effectively, the dev-centric software engineering (aka “data engineering”) approach and the statistics detour of the past two centuries – are losing steam and lacked sufficient depth to begin with. Machine learning in the 1980s meant something much broader than what gets represented by the current crop of analytics vendors; check out my preso for more details. To cut to the chase, also check an excellent talk The Thorn in the Side of Big Data: too few artists by Christopher Ré @Stanford. See a related article I’ll Be Back: The Return of Artificial Intelligence by Jack Clark @BloombergBusiness.


Stanford Y2E2 at sunset

I have a hunch that cloud-based notebooks will eat the lunch of oh-so-many dev-centric approaches and second-generation BI tools. That strips away from the intrinsic value of Hortonworks, Cloudera, etc. Meanwhile it pushes value toward those firms which are closest to domain experts, with key examples such as Enlitic, Idibon, Oculus Info, Spaceknow, etc.

The fourth category has a large market in industry in general. In my opinion, going forward its upside will be realized less so among the “data-centric” usual suspects of ad tech, fin tech, e-commerce, social networks, security… rather more so within the more traditional sectors of energy, transportation, manufacturing, agriculture, etc. Sensor data is a major driver, whether we are talking about embedded sensors or layers of remote sensing or for that matter the volumes of data in genomics work. These use cases tend toward streaming. Fine-grained resource management in clusters is core to this: not so much due to the data rates as it is due to needs for elastic computing capacity and service architectures – in other words, latency and robustness become key. Streaming applications have lots of moving parts and represent a hard problem in computer science in general. On the one hand, the organizational costs of using a YARN cluster to address those kinds of needs proves to be rather upside down, while on the other hand we see a rise in Mesos deployments, e.g., Virdata, Atigeo, Stratio, etc.

My hunch is that the emerging stack for sophisticated analytics and optimization needs will look significantly less like Cloudera or Hortonworks, and more like a integration of...

Typesafe is another vendor that is clearly addressing this demand. However, that speaks to the infrastructure not the science, and this is where the focus on tensors comes back into the picture…

Within the 2–3 year horizon, I expect to see reasonably good open source projects for cost-effective and scalable methods for low-rank tensor factorization. It’s likely this will involve some probabilistic techniques and lead toward online algorithms, i.e., for streaming. So far there haven’t been good off-the-shelf solutions for tensor factorization. However, a general case approach that could scale-out on commodity hardware would be a significant game-changer, with the potential to sublate a wide range of contemporary work in algorithms.

Within a similar timeline, I expect to see relatively dramatic improvements in networking technology, i.e., within the datacenter. Taken together those two events would signal the availability of relatively more general purpose solutions in contrast to the many one-offs in analytics that are currently bread-and-butter for Hadoop app developers. It could also erode the valuation for the many machine learning library vendors. Consequently, I’m watching this area closely as the sea change evolves. 

My prediction about Hadoop was on target, so let’s see how this new prediction unfolds.

Spark

We’ve had the Apache Spark developer certificate available online for several weeks now. Congrads to the recipient of certificate number 1.1.0 - 0001, François Garrilot @Typesafe. While I cannot release exact numbers, the success rate for people taking the exam is in the mid 90’s percent. It pays to have hands-on experience developing Spark apps, and this talk provides some great test prep examples. We’ll work toward certifications that are more specialized toward systems engineering and data science.


First Spark certificate goes to François Garrilot!

Recently, Reynold Xin presented about the new DataFrames support in Spark, bringing parity with similar abstractions in Python and R. This capability will be introduced but disabled by default in Spark 1.3, but will become center-stage in later releases. In terms of workflows, it represents a higher-level abstraction than RDDs; however, there are still RDDs underneath and many applications will continue to focus at that layer. Meanwhile, Matei’s thesis has been translated into Chinese. Hopefully that represents the beginning of trend.

Also check out the events worldwide listings and archived talks on the YouTube channel for Apache Spark.

Workplace

So much effort these days seems to be spent on achieving #Inbox40 … I have a hunch that use of email for business must be rethought. Soon. And perhaps abandoned? I am not convinced that productivity tools such Yammer, Asana, Slack, etc., provide any long-term solutions, since they still tend to focus people too much on screens and keyboards.


Pescadero Beach, office for an afternoon on the way from our company retreat

FWIW, among my daughters’ peer group, they are way more Internet-savvy than #millenials and have already dumped email as #deadmedia … They use Instagram, Minecraft, and Skype as collaboration tools – each of which is at least partly owned by MSFT, for those who are keeping track. However, they concede that they’d likely use Twitter for business if they needed it. Consequently, I greatly appreciate when people use my public timeline on Twitter to communicate. At this point, I delete most private messages aside from Gmail: Twitter DMs, LinkedIn mail, etc., and Gmail messages are N-deep before they will get read.

Just Enough Math

Apparently the Foobartendr drink-by-drone-delivery service in Just Enough Math wasn’t so cray-cray after all ;) Recently the Washington Post reported about a restaurant delivering drinks via drones indoors.

Another interesting bit of tech news is in Quantum Information Processing: Are We There Yet? by Daniel Lidar @USC: niobium processors, Chimera graphs, and much more fun. To wit, this video discusses how to solve Ising Hamiltonians with quantum annealing, i.e., for complex graph problems. Gosh, wonder if that could be handy for tensor factorization? Check around the 36:48 mark, where Prof. Lidar discusses how ground state success probability distributions for DWave are inconsistent with thermal annealer (classical / unimodal) results, but consistent with simulated quantum annealer (bimodal). As far as I can follow the discussion, this rules out classical models, but is not definitive proof yet. Also, how well will it scale?

Upcoming Events

Many interesting conferences and other events are planned for the months ahead… Please check the http://goo.gl/2YqJZK listings. In particular, mark your calendars for:
Meanwhile we’re busy preparing for Spark Summit East next month in NYC on Mar 18–19. Please join us, and to help with that here’s a 20% discount code SSPACO20 for registration.

Also, make plans for MesosCon 2015, Aug 20–21 in Seattle.

Misc.

Just under the wire: for what it’s worth, I barely squeaked into the Top 30 People in Big Data and Analytics and also recently joined the academic advisory board for the GalvanizeU graduate program in data science. Grateful for both of those.

Whenever I go to write a newsletter, I’m concerned that there won’t be enough content collected yet. Invariably, there are too many links to share. Here are some that caught my attention recently…

The Africa soil map shows the changing nature of soil across the continent. as “an essential reference to a non-renewable resource that is fundamental for life on this planet.” A vital lesson to all, for there are no jobs on a dead planet. Establishing a bar here, I wish we had comparable analysis for North America.

Perhaps one of the more jaw-dropping research results recently: photonic radiative cooling by Shanhui Fan, et al., @Stanford. More than simply an enormous increase in the capability for buildings to reflect sunlight efficiently, this provides a way to beam internal heat out into space without warming the atmosphere: “What we’ve done is to create a way that should allow us to use the coldness of the universe as a heat sink during the day.”

Another interesting development is the US Digital Service: “The United States Digital Service is transforming how the federal government works for the American people. And we need you.” That emerges along with DJ Patil becoming US Chief Data Scientist.

Following that, I’ll leave you with something fun and something epic. First, a limerick detector, based on the GitHub repo Nantucket. Second, words of wisdom from Vint Cerf: Forgotten Century.


That's the update for now. See you in NYC, Boulder, São Paulo, Boston, London, A Coruña, and Chicago on the event horizon!

2014-12-29

Newsletter Updates for December 2014

Chicago, Boulder, NYC, DC, SF, Stanford, London, Stockholm, Madrid, Barcelona, Amsterdam, Dulles, Baltimore, LA. The range of speaking events and business travel over the past quarter almost bewilders, but I’m grateful to get to meet many interesting people and learn about new projects. 

Also feeling grateful to enjoy some quiet time at home with family over the holidays, and I wish very happy holidays to you and yours.

Conference Summaries

Strata NY set a new record with about 450 people attending Spark Camp. There was a spare room, plus an hour break in the fray, so we held an impromptu “Ask Us Anything” about Spark – that has turned into a new kind of open source ritual at Strata confs, especially for handling the more advanced audience questions. Also, Bloomberg kindly hosted a large Spark Committer Night meetup event, their largest to-date.

Manhattan, from NY Water Taxi at Port Imperial
Throughout many conferences and meetup events over the past few months, one demo in particular stood out. David Jonker and Rob Harper from Oculus Info in Toronto gave a talk about Aperture Tiles at Strata NY. Last talk of the show, and quite arguably the best. This open source framework, partly built atop Spark, provides interactive data exploration with continuous zooming on large scale datasets. Highly recommended.

The week after Strata conf in NYC, some of our team found our way slightly south to the University of Maryland, where we got to teach alongside the renowned Jimmy Lin. The week included a Spark Tutorial on campus, plus the initial meeting of the Apache Spark Maryland meetup. Much fun, and we look forward to returning to UMD again soon.

Arriving back to the Bay Area just in time, I caught the launch of the new GalvanizeU program in downtown SF. One challenge that particular evening was getting scheduled to speak head-to-head with the final game in the World Series. That keynote, Data Science in Future Tense, examined some of the near-past and near-future of the field – hopefully indicating some non-intuitive directions. 

GalvanizeU is located next to the Transbay Center, just a few blocks away from the new Databricks office. They provide a hands-on graduate program in Data Science, in an urban setting and working closely with industry partners. Galvanize started in Boulder and is also expanding soon into Seattle. We’re thrilled about our new neighbors.

Home just long enough to take the kiddos trick-or-treating and attend GCPLive, then on to Europe… During a brief visit in the UK, I got to present about the latest in Spark Streaming at the London Spark Meetup: Tiny Batches in the wine (a callback to Don Ho, for those who were born more recently – ideal for getting your luau on). Then on to Stockholm with gracious hosting by Spotify, Ericsson, and SICS.

Good times @ Big Data Spain, Madrid

Madrid came next, for the annual Big Data Spain conf. Noticing a joke painted on the side of a jet at the airport, I had a hunch immediately that Madrid would be lots of fun. I was not disappointed. Our hosts at Paradigma Tecnólogico and Stratio presented an amazing conference, one of my favorites in a long, long time. I was fortunate to give a keynote talk, alongside many other excellent talks, such as from friends at Cloudera and Google BigQuery. I highly recommend Big Data Spain. More about Stratio in a bit…

The beach at El Poblenou, Barcelona
Taking a train from Madrid to Barcelona, admittedly I was missing the former, but Barcelona is a wonderful place. Imagine yourself in Santa Barbara, except that the city is 50 times larger, thousands of years older, and packed full o’ amazing culture. Strata EU was located at a conference center right next to the beach. We held the first official Spark developer certificate exam, plus a large Spark Camp event (25% of the conference attended), a meetup at UPC, and a second iteration of our “Ask Us Anything” about Spark.

Locavore feasting in Catalunya
Business travel Spark-style does not allow much downtime. Effectively one day off during two full weeks in EU. Fortunately that just happened to be during a weekend in Barcelona, the day after Strata concluded. I rented an Airbnb condo near the beach in El Poblenou, then wandered busy Rambla markets, through the crowd surrounding a busker string trio, gathering items to make a small feast. Only in Catalunya.

Amstel River in Amsterdam
A quick stop in Amsterdam, with a very fun talk hosted at eBay with hours of Q&A, then back home. Long enough for a family Thanksgiving feast, then off to DC, Baltimore, and LA. Excellent events and good friends met along the way, particularly the Los Angeles Apache Spark meetup hosted by Rubicon Project. Much appreciated.

Spark

The curiously named Likelihood T. Prior noted on Twitter: Spark spark spark spark, spark spark spark spark. #Strataconf synopsis complete. Some went as far as to begin calling “Strata + Hadoop World” by a new name, “Strata + Spark World”. I like the sound of that.

To help keep track of this rocket ride, I’ve begun curating an ongoing list http://goo.gl/2YqJZK of the talks, workshops, etc., related to Spark worldwide. Please let me know if you have events to add.
Speaking of events, recently we began to increase the cadence for Bay Area Spark meetup events. These talks get live-streamed, with the archives published on the Apache Spark channel on YouTube. Databricks also recently announced Spark Packages a community index of packages. The site had to be moved shortly after its launch, due to overwhelming popularity. Good stuff on both the video channel and package repo.

So much news about Spark has happened in the past few months. I’d like to summarize with a few gems collected along the way…
Not least of these items, the Databricks team broke YHOO’s previous world record for the Daytona GraySort contest. That tied for the 100 PB sort on AWS, using 1/10 the number of servers and running 3x faster than YHOO Hadoop clusters. #justsayin

MOOCs

Part of my job involves the curriculum for Spark instruction. Our big news recently is that edX and the University of California will be offering two new MOOCs about Spark, sponsored by Databricks.

The first is Introduction to Big Data with Apache Spark by Prof. Anthony Joseph at UC Berkeley. This comprehensive introduction to Spark, as well as Big Data, is based entirely on Python programming and aimed at developing Data Science skills. This course begins on 2015–02–23.

The second is Scalable Machine Learning by Prof. Ameet Talwalkar at UCLA. This hands-on course focuses on distributed machine learning at scale, based on examples using open data, also in Python. This course begins on 2015–04–14.

Note that some taking Spark MOOCs will have the option to use Databricks Cloud free student accounts. Similarly, we will be integrating use of DBC free accounts into our other Spark training events.

Workplace

Several years ago, I was fortunate to work for a CEO who understood how to leverage a distributed workplace. I studied the management practices involved, and in particular have grown to appreciate ROWE greatly. These practices seem all too rare among early-stage tech start-ups Silicon Valley. However, a few tech firms (DataStax and Typesafe come to mind) have embraced distributed workplace models. Frankly, correlations between effective approaches to gender equality and practices such as ROWE should be on every VC’s radar.

With respect to workplace practices – effective or otherwise – two recent articles caught my attention:
Great words of wisdom about two of the worst anti-patterns for successful tech organizations. The most telling part is the “canary in a coal mine” effect: to watch and see who becomes the most offended by these points. Egregious (sometimes outright hostile) use of email, chat, meetings, etc., and the fallacy of “crunch mode” stand as two of my top determinants for evaluating a company. Right alongside we provide free snacks and meals vs. we offer reasonable health care plans – which somehow turn out to be at odds in far too many start-ups.

BTW, really looking forward to catching Chad speak at GOTO Chicago next May.

Just Enough Math

The Just Enough Math material continues to evolve… Allen and I gave a tutorial at Strata NY, working closely with O’Reilly Media to export content to IPython Notebook within a Docker container for participants to run in the cloud. Rackspace provided the hosting, which in turn was an alpha test for their Nature magazine IPython interactive demo. Welcome to the future of publishing.

Andrew Odewahn and I entered a version of this for the Boston instance of Docker Global Hack Day #2 – frankly, Andrew did like 99.9999% of the work on that one :) Meanwhile, speaking of the future of publishing, JEM provides an example in the new Publishing Workflows for Jupyter by Andrew Odewahn, Kyle Kelley, Rune Madsen.

Beyond publishing, we do have some math to suggest… Two papers caught my attention recently:
Oh, and riffing off the “Quantum Algorithms on the Moon” meme from JEM, note that NASA, Google and USRA establish Quantum Computing Research Collaboration such that 20% of computing time will be provided to the university community. In case you have some large data set that’s just screaming to get crunched on a D-Wave. Like you do.

Mesos


Other big news was the Google Cloud Platform Live conference in SF on Nov 1. The message from #GCPLive was largely about containers… in short, the notion of The datacenter IS the computer going mainstream. To paraphrase one comment during the conf: “Customers get locked into host-based patterns, so they struggle with intertwined systems.” Well said. Definitely looking forward to the new GKE service based on Kubernetes.

Other big news was awaiting in London. Namely, the team behind Weave. Recall that the JEM tutorial had been an alpha test for the IPython + Docker + Rackspace + Nature magazine thing? We learned a truism the hard way, with minutes to go before the event started: Docker does little to resolve crucial issues outside of the containers. Enter Weave, handling difficult matters outside the container, such as networking and crypto. Check their blog for tasty insights, e.g., Automated provisioning of multi-cloud weave network with Terraform. Highly recommended.

Speaking of Docker, I really enjoyed this talk by Adrian Cockcroft @DockerCon: State of the Art in Microservices. Especially slides #8–19, product development process.

Speaking of Microservices, here’s a good overview: The Strengths and Weaknesses of Microservices by Abel Avram on InfoQ.

Ag+Data

Continuing on the Ag+Data front, check out the excellent article GeoTrellis Adapts to Climate Change and Spark about how Climate Change analytics drove Spark adoption at Azavea. They integrated Spark and Accumulo to support fast computation of climate impact metrics for DoE, which should be included in the 0.10 release of GeoTrellis.

NYT ran an interactive analysis/visualization, Flooding Risk From Climate Change, Country by Country, which perhaps helps explain Silicon Valley rumors about Google building ferry ports at corporate campuses along SF Bay.

I’m a big fan of Danielle Nierenberg @FoodTank in Chicago. A recent article, How Vegetables Can Save the World, is brief, accessible, and quite to the point. More of that on FoodTank.
Meanwhile, considering the many challenges ahead in Ag worldwide, I’m curious whether some programmable matter could become useful on farms to leverage data? Sort of an asymptote for IoT.

Upcoming Events

Many interesting conferences and other events are planned for the months ahead. Please do check the http://goo.gl/2YqJZK listings. In particular, mark your calendars for:
O'Reilly studio in Sebastopol, for new "Intro Spark" video

Misc.

I’ll leave you with something fun and something epic.

First, the fun – though it’s quite epic in a way: LumiGeek. We make Arduino shields for LEDs, audio-reactive drivers, and custom solutions for architectural and artistic endeavors. Check their installation at the new Galvanize Cafe in SF, and look about carefully for a subtle case of anamorphosis.

Second, the epic – if you haven’t seen it yet, it’s well worth four gorgeous minutes of video: Wanderers by Erik Wernquist, narrated by Carl Sagan. Money quote @1:45: “Herman Melville in Moby Dick spoke for wanderers in all epochs and meridians…”

That's the update for now. See you in Austin, San Jose, and NYC on the event horizon!