Showing posts with label startups. Show all posts
Showing posts with label startups. Show all posts

2015-03-07

SV Synopsis: Fundamentalism in Technology

I am grateful for perspectives gained because our family lives in Silicon Valley. Many options here to work at novel ventures, and on fascinating projects… Opportunities to drop by Stanford or Berkeley for some remarkable guest lecture by a visiting expert… The wonders of an almost perpetual Maker Faire as one walks through the neighborhood on any given evening… Tech camps that our daughters can attend locally, as they wish… And, generally speaking, the lack of any real need to engage in ridiculous commutes

As an open source evangelist and as an investor, I've felt grateful to learn from a veritable parade of interesting projects. However, I am troubled by the incidence of a particular problem. Far too often one runs headlong into what I could characterize as a close approximation of cocaine-fueled misogynistic narcissism. The condition is subtle, but systemic here. Even recently, I have witnessed this up close – along with the regrettably pervasive and predictable non-reactions to it. Increasingly, zero tolerance appears to be the only effective response. Or perhaps the tech industry percolates out elsewhere, far from SF and its inertia?

Without mentioning names, two well-known billionaire-club investors in Silicon Valley personify this character sketch. Evidence of panspermia ad absurdum festers in the "cultures" that they promote. Personal jihads seemingly to self-perpetuate their fundamentalist ideals.

A nagging question lingers… Why work alongside an ilk of people with whom I would never encourage my daughters to mingle? Granted, I believe quite strongly in the need to talk with just about everyone, to keep dialogue open, to reject the notion of "enemy". Even so, there are absolutes. Practical realities of livelihood aside, as a parent what kind of examples do my professional actions and affiliations set?

In addition, a question that investors ask over and over when considering whether to fund a new company is "Will the team scale?" Any measure of the toxins described above almost guarantees that the answer will in practice be "No."


That represents a dirty little secret. There is an amazing level of demand for tech talent. It's not exactly because these companies are raging commercial successes; most early-stage ventures by definition are not. It's because few people who are capable of making good judgements are willing to compromise their futures to work for ineffective caricatures. Many start-ups encounter difficulties in scaling their team. Or – more likely over time – they encounter high attrition rates.

While I have in the past focused for several years on the same project, lately I don't stay long in most early-stage firms, generally moving on after an organization demonstrates its nature. To paraphrase Lady Grantham from Downton Abbey, there is a point at which malice ceases to be amusing. On the one hand, that's a terrible way to leverage stock option packages. On the other hand, arguably I have pursued a portfolio career strategy. That approach has helped me build an amazing network. Long-term benefits of my network have far surpassed the potential upside of my aggregate stock options. Therein dwells an important lesson about Silicon Valley.

2015-02-28

Newsletter Updates for February 2015

Not so much travel recently – Austin was my only trip this quarter so far. We’ve been heads-down reworking instructional materials to highlight what you can do with cloud-based notebooks. To learn more about that, check out the new Databricks newsletter.

Snow near Cold Springs, California

Meanwhile, my family gets to enjoy some time this weekend in a cabin near Yosemite, during an increasingly rare event here: lots of snow! Recommend: we always try to drop by our favorite mile-high restaurant, Mia’s, for excellent Italian cooking in the mountains and even homemade limoncello.

Strata

Of course, one of the other big reasons for keeping close to home lately was our biggest event of the year, Strata + Hadoop World in San Jose. Here’s a link for the published speaker slides and videos, along with an excellent summary of the Hardcore Data Science day by Ben Lorica.

About 325 people attended our Spark Camp tutorial. Oddly enough, that’s the same ratio of total conference attendees that we had at Spark Camp in NYC last fall. I also got to host the new Spark in Action track. One eye-opener in our track was the Tencent talk, where LianHui Wang presented about their experiences running an 8000 node Spark cluster in production. So much for FUD claims that Spark doesn’t scale ;) When asked how Tencent can build substantially larger clusters than what YHOO has reported, LianHui replied wryly, “They do not speak Chinese.”

StackOverflow analysis of Spark by Donnie Berkholz @RedMonk

One of the other Strata talks that I really wanted to catch: Tensor Methods for Large-scale Unsupervised Learning: Applications to Topic and Community Modeling by Animashree Anandkumar @UC Irvine. For more details, check out her video

In particular note the experimental results at the 42:46 mark, along with slides for a related talk. There is even more background in the recent papers: Guaranteed Non-Orthogonal Tensor Decomposition via Alternating Rank–1 Updates; and Tensor decompositions for learning latent variable models.

The gist of this effort is about using graph moments, assuming priors which then help make tensor decomposition tractable. This material will flex your advanced math agility as it flies through linear algebra, graph theory, statistics, and optimization for some startling implications. While the immediate research is about latent variables for community detection (think: Facebook) these techniques have implications on a much broader range of industry optimization problems. Note that the outcomes are in contrast to work by Jure Leskovec, et al., @Stanford. Another excellent Spark-related talk at Strata that referenced work with tensors was Hadoop as a Platform for Genomics by Allen Day @MapR .

Looking Ahead

Why tensors? Recall from 18 months ago, “I give Hadoop three years before it gets displaced.” At the time that prediction drew some flack. Now that we’re halfway to the predicted time, note that during the past three Strata + Hadoop World conferences there have been numerous remarks to rename it Strata + Spark World. However, the general insight drives a bit deeper…

My question here is, “What is the business case for developing custom apps atop a Hadoop platform?” When I examine industry use cases for Big Data frameworks, there are a few general categories:
  1. ETL
  2. data warehouse replacements
  3. data exploration and reporting
  4. analytics in depth, leading toward streaming
The first category is relatively well-understood, leading toward general purpose solutions. On the start-up side of the spectrum there are great solutions emerging such as ETLeapAlation, and arguably examples such as Epic in medical data exchange. On the established side of Enterprise IT, incumbents such as Informatica have been aggressively partnering and expanding the scope of their integration. That begs the question of whether firms would continue to build rather than buy?

The second and third categories are the devil-you-know, as continuations of DW and BI respectively. SiliconAngle had a good article recently along these lines, The cheat sheet to following Big Data’s money trail by Suzanne Kattau.

My hunch is that in terms of the second category, Cloudera, Hortonworks, etc., will be forced to pivot toward vertical applications sooner than later to sustain their growth, and will likely buy up smaller analytics vendors along the way. That puts them on a collision course with incumbents Oracle, IBM, Teradata, SAS, etc., where both ends of the spectrum race toward resembling each other. In other words, the DW king is dead, long live the DW king. Expect either some contractions or M&A activity as a result. Not much news there.

The third category, effectively a BI displacement, gets a bit more interesting. I gave a keynote talk at Data Day Texas in Austin in January, A New Year in Data Science: ML Unpaused. The gist is that two aspects of the BI displacement – effectively, the dev-centric software engineering (aka “data engineering”) approach and the statistics detour of the past two centuries – are losing steam and lacked sufficient depth to begin with. Machine learning in the 1980s meant something much broader than what gets represented by the current crop of analytics vendors; check out my preso for more details. To cut to the chase, also check an excellent talk The Thorn in the Side of Big Data: too few artists by Christopher Ré @Stanford. See a related article I’ll Be Back: The Return of Artificial Intelligence by Jack Clark @BloombergBusiness.


Stanford Y2E2 at sunset

I have a hunch that cloud-based notebooks will eat the lunch of oh-so-many dev-centric approaches and second-generation BI tools. That strips away from the intrinsic value of Hortonworks, Cloudera, etc. Meanwhile it pushes value toward those firms which are closest to domain experts, with key examples such as EnliticIdibonOculus InfoSpaceknow, etc.

The fourth category has a large market in industry in general. In my opinion, going forward its upside will be realized less so among the “data-centric” usual suspects of ad techfin teche-commercesocial networkssecurity… rather more so within the more traditional sectors of energy, transportation, manufacturing, agriculture, etc. Sensor data is a major driver, whether we are talking about embedded sensors or layers of remote sensing or for that matter the volumes of data in genomics work. These use cases tend toward streaming. Fine-grained resource management in clusters is core to this: not so much due to the data rates as it is due to needs for elastic computing capacity and service architectures – in other words, latency and robustness become key. Streaming applications have lots of moving parts and represent a hard problem in computer science in general. On the one hand, the organizational costs of using a YARN cluster to address those kinds of needs proves to be rather upside down, while on the other hand we see a rise in Mesos deployments, e.g., VirdataAtigeoStratio, etc.

My hunch is that the emerging stack for sophisticated analytics and optimization needs will look significantly less like Cloudera or Hortonworks, and more like a integration of...

Typesafe is another vendor that is clearly addressing this demand. However, that speaks to the infrastructure not the science, and this is where the focus on tensors comes back into the picture…

Within the 2–3 year horizon, I expect to see reasonably good open source projects for cost-effective and scalable methods for low-rank tensor factorization. It’s likely this will involve some probabilistic techniques and lead toward online algorithms, i.e., for streaming. So far there haven’t been good off-the-shelf solutions for tensor factorization. However, a general case approach that could scale-out on commodity hardware would be a significant game-changer, with the potential to sublate a wide range of contemporary work in algorithms.

Within a similar timeline, I expect to see relatively dramatic improvements in networking technology, i.e., within the datacenter. Taken together those two events would signal the availability of relatively more general purpose solutions in contrast to the many one-offs in analytics that are currently bread-and-butter for Hadoop app developers. It could also erode the valuation for the many machine learning library vendors. Consequently, I’m watching this area closely as the sea change evolves. 

My prediction about Hadoop was on target, so let’s see how this new prediction unfolds.

Spark

We’ve had the Apache Spark developer certificate available online for several weeks now. Congrads to the recipient of certificate number 1.1.0 - 0001François Garrilot @Typesafe. While I cannot release exact numbers, the success rate for people taking the exam is in the mid 90’s percent. It pays to have hands-on experience developing Spark apps, and this talk provides some great test prep examples. We’ll work toward certifications that are more specialized toward systems engineering and data science.


First Spark certificate goes to François Garrilot!

Recently, Reynold Xin presented about the new DataFrames support in Spark, bringing parity with similar abstractions in Python and R. This capability will be introduced but disabled by default in Spark 1.3, but will become center-stage in later releases. In terms of workflows, it represents a higher-level abstraction than RDDs; however, there are still RDDs underneath and many applications will continue to focus at that layer. Meanwhile, Matei’s thesis has been translated into Chinese. Hopefully that represents the beginning of trend.

Also check out the events worldwide listings and archived talks on the YouTube channel for Apache Spark.

Workplace

So much effort these days seems to be spent on achieving #Inbox40 … I have a hunch that use of email for business must be rethought. Soon. And perhaps abandoned? I am not convinced that productivity tools such YammerAsanaSlack, etc., provide any long-term solutions, since they still tend to focus people too much on screens and keyboards.


Pescadero Beach, office for an afternoon on the way from our company retreat

FWIW, among my daughters’ peer group, they are way more Internet-savvy than #millenials and have already dumped email as #deadmedia … They use InstagramMinecraft, and Skype as collaboration tools – each of which is at least partly owned by MSFT, for those who are keeping track. However, they concede that they’d likely use Twitter for business if they needed it. Consequently, I greatly appreciate when people use my public timeline on Twitter to communicate. At this point, I delete most private messages aside from Gmail: Twitter DMs, LinkedIn mail, etc., and Gmail messages are N-deep before they will get read.

Just Enough Math

Apparently the Foobartendr drink-by-drone-delivery service in Just Enough Math wasn’t so cray-cray after all ;) Recently the Washington Post reported about a restaurant delivering drinks via drones indoors.

Another interesting bit of tech news is in Quantum Information Processing: Are We There Yet? by Daniel Lidar @USC: niobium processors, Chimera graphs, and much more fun. To wit, this video discusses how to solve Ising Hamiltonians with quantum annealing, i.e., for complex graph problems. Gosh, wonder if that could be handy for tensor factorization? Check around the 36:48 mark, where Prof. Lidar discusses how ground state success probability distributions for DWave are inconsistent with thermal annealer (classical / unimodal) results, but consistent with simulated quantum annealer (bimodal). As far as I can follow the discussion, this rules out classical models, but is not definitive proof yet. Also, how well will it scale?

Upcoming Events

Many interesting conferences and other events are planned for the months ahead… Please check the http://goo.gl/2YqJZK listings. In particular, mark your calendars for:
Meanwhile we’re busy preparing for Spark Summit East next month in NYC on Mar 18–19. Please join us, and to help with that here’s a 20% discount code SSPACO20 for registration.

Also, make plans for MesosCon 2015, Aug 20–21 in Seattle.

Misc.

Just under the wire: for what it’s worth, I barely squeaked into the Top 30 People in Big Data and Analytics and also recently joined the academic advisory board for the GalvanizeU graduate program in data science. Grateful for both of those.

Whenever I go to write a newsletter, I’m concerned that there won’t be enough content collected yet. Invariably, there are too many links to share. Here are some that caught my attention recently…

The Africa soil map shows the changing nature of soil across the continent. as “an essential reference to a non-renewable resource that is fundamental for life on this planet.” A vital lesson to all, for there are no jobs on a dead planet. Establishing a bar here, I wish we had comparable analysis for North America.

Perhaps one of the more jaw-dropping research results recently: photonic radiative cooling by Shanhui Fan, et al., @Stanford. More than simply an enormous increase in the capability for buildings to reflect sunlight efficiently, this provides a way to beam internal heat out into space without warming the atmosphere: “What we’ve done is to create a way that should allow us to use the coldness of the universe as a heat sink during the day.”

Another interesting development is the US Digital Service: “The United States Digital Service is transforming how the federal government works for the American people. And we need you.” That emerges along with DJ Patil becoming US Chief Data Scientist.

Following that, I’ll leave you with something fun and something epic. First, a limerick detector, based on the GitHub repo Nantucket. Second, words of wisdom from Vint Cerf: Forgotten Century.


That's the update for now. See you in NYC, Boulder, São Paulo, Boston, London, A Coruña, and Chicago on the event horizon!

2014-12-29

Newsletter Updates for December 2014

Chicago, Boulder, NYC, DC, SF, Stanford, London, Stockholm, Madrid, Barcelona, Amsterdam, Dulles, Baltimore, LA. The range of speaking events and business travel over the past quarter almost bewilders, but I’m grateful to get to meet many interesting people and learn about new projects. 

Also feeling grateful to enjoy some quiet time at home with family over the holidays, and I wish very happy holidays to you and yours.

Conference Summaries

Strata NY set a new record with about 450 people attending Spark Camp. There was a spare room, plus an hour break in the fray, so we held an impromptu “Ask Us Anything” about Spark – that has turned into a new kind of open source ritual at Strata confs, especially for handling the more advanced audience questions. Also, Bloomberg kindly hosted a large Spark Committer Night meetup event, their largest to-date.

Manhattan, from NY Water Taxi at Port Imperial
Throughout many conferences and meetup events over the past few months, one demo in particular stood out. David Jonker and Rob Harper from Oculus Info in Toronto gave a talk about Aperture Tiles at Strata NY. Last talk of the show, and quite arguably the best. This open source framework, partly built atop Spark, provides interactive data exploration with continuous zooming on large scale datasets. Highly recommended.

The week after Strata conf in NYC, some of our team found our way slightly south to the University of Maryland, where we got to teach alongside the renowned Jimmy Lin. The week included a Spark Tutorial on campus, plus the initial meeting of the Apache Spark Maryland meetup. Much fun, and we look forward to returning to UMD again soon.

Arriving back to the Bay Area just in time, I caught the launch of the new GalvanizeU program in downtown SF. One challenge that particular evening was getting scheduled to speak head-to-head with the final game in the World Series. That keynote, Data Science in Future Tense, examined some of the near-past and near-future of the field – hopefully indicating some non-intuitive directions. 

GalvanizeU is located next to the Transbay Center, just a few blocks away from the new Databricks office. They provide a hands-on graduate program in Data Science, in an urban setting and working closely with industry partners. Galvanize started in Boulder and is also expanding soon into Seattle. We’re thrilled about our new neighbors.

Home just long enough to take the kiddos trick-or-treating and attend GCPLive, then on to Europe… During a brief visit in the UK, I got to present about the latest in Spark Streaming at the London Spark Meetup: Tiny Batches in the wine (a callback to Don Ho, for those who were born more recently – ideal for getting your luau on). Then on to Stockholm with gracious hosting by Spotify, Ericsson, and SICS.

Good times @ Big Data Spain, Madrid

Madrid came next, for the annual Big Data Spain conf. Noticing a joke painted on the side of a jet at the airport, I had a hunch immediately that Madrid would be lots of fun. I was not disappointed. Our hosts at Paradigma Tecnólogico and Stratio presented an amazing conference, one of my favorites in a long, long time. I was fortunate to give a keynote talk, alongside many other excellent talks, such as from friends at Cloudera and Google BigQuery. I highly recommend Big Data Spain. More about Stratio in a bit…

The beach at El Poblenou, Barcelona
Taking a train from Madrid to Barcelona, admittedly I was missing the former, but Barcelona is a wonderful place. Imagine yourself in Santa Barbara, except that the city is 50 times larger, thousands of years older, and packed full o’ amazing culture. Strata EU was located at a conference center right next to the beach. We held the first official Spark developer certificate exam, plus a large Spark Camp event (25% of the conference attended), a meetup at UPC, and a second iteration of our “Ask Us Anything” about Spark.

Locavore feasting in Catalunya
Business travel Spark-style does not allow much downtime. Effectively one day off during two full weeks in EU. Fortunately that just happened to be during a weekend in Barcelona, the day after Strata concluded. I rented an Airbnb condo near the beach in El Poblenou, then wandered busy Rambla markets, through the crowd surrounding a busker string trio, gathering items to make a small feast. Only in Catalunya.

Amstel River in Amsterdam
A quick stop in Amsterdam, with a very fun talk hosted at eBay with hours of Q&A, then back home. Long enough for a family Thanksgiving feast, then off to DC, Baltimore, and LA. Excellent events and good friends met along the way, particularly the Los Angeles Apache Spark meetup hosted by Rubicon Project. Much appreciated.

Spark

The curiously named Likelihood T. Prior noted on TwitterSpark spark spark spark, spark spark spark spark. #Strataconf synopsis complete. Some went as far as to begin calling “Strata + Hadoop World” by a new name, “Strata + Spark World”. I like the sound of that.

To help keep track of this rocket ride, I’ve begun curating an ongoing list http://goo.gl/2YqJZK of the talks, workshops, etc., related to Spark worldwide. Please let me know if you have events to add.
Speaking of events, recently we began to increase the cadence for Bay Area Spark meetup events. These talks get live-streamed, with the archives published on the Apache Spark channel on YouTube. Databricks also recently announced Spark Packages a community index of packages. The site had to be moved shortly after its launch, due to overwhelming popularity. Good stuff on both the video channel and package repo.

So much news about Spark has happened in the past few months. I’d like to summarize with a few gems collected along the way…
Not least of these items, the Databricks team broke YHOO’s previous world record for the Daytona GraySort contest. That tied for the 100 PB sort on AWS, using 1/10 the number of servers and running 3x faster than YHOO Hadoop clusters. #justsayin

MOOCs

Part of my job involves the curriculum for Spark instruction. Our big news recently is that edX and the University of California will be offering two new MOOCs about Spark, sponsored by Databricks.

The first is Introduction to Big Data with Apache Spark by Prof. Anthony Joseph at UC Berkeley. This comprehensive introduction to Spark, as well as Big Data, is based entirely on Python programming and aimed at developing Data Science skills. This course begins on 2015–02–23.

The second is Scalable Machine Learning by Prof. Ameet Talwalkar at UCLA. This hands-on course focuses on distributed machine learning at scale, based on examples using open data, also in Python. This course begins on 2015–04–14.

Note that some taking Spark MOOCs will have the option to use Databricks Cloud free student accounts. Similarly, we will be integrating use of DBC free accounts into our other Spark training events.

Workplace

Several years ago, I was fortunate to work for a CEO who understood how to leverage a distributed workplace. I studied the management practices involved, and in particular have grown to appreciate ROWE greatly. These practices seem all too rare among early-stage tech start-ups Silicon Valley. However, a few tech firms (DataStax and Typesafe come to mind) have embraced distributed workplace models. Frankly, correlations between effective approaches to gender equality and practices such as ROWE should be on every VC’s radar.

With respect to workplace practices – effective or otherwise – two recent articles caught my attention:
Great words of wisdom about two of the worst anti-patterns for successful tech organizations. The most telling part is the “canary in a coal mine” effect: to watch and see who becomes the most offended by these points. Egregious (sometimes outright hostile) use of email, chat, meetings, etc., and the fallacy of “crunch mode” stand as two of my top determinants for evaluating a company. Right alongside we provide free snacks and meals vs. we offer reasonable health care plans – which somehow turn out to be at odds in far too many start-ups.

BTW, really looking forward to catching Chad speak at GOTO Chicago next May.

Just Enough Math

The Just Enough Math material continues to evolve… Allen and I gave a tutorial at Strata NY, working closely with O’Reilly Media to export content to IPython Notebook within a Docker container for participants to run in the cloud. Rackspace provided the hosting, which in turn was an alpha test for their Nature magazine IPython interactive demo. Welcome to the future of publishing.

Andrew Odewahn and I entered a version of this for the Boston instance of Docker Global Hack Day #2 – frankly, Andrew did like 99.9999% of the work on that one :) Meanwhile, speaking of the future of publishing, JEM provides an example in the new Publishing Workflows for Jupyter by Andrew OdewahnKyle KelleyRune Madsen.

Beyond publishing, we do have some math to suggest… Two papers caught my attention recently:
Oh, and riffing off the “Quantum Algorithms on the Moon” meme from JEM, note that NASA, Google and USRA establish Quantum Computing Research Collaboration such that 20% of computing time will be provided to the university community. In case you have some large data set that’s just screaming to get crunched on a D-Wave. Like you do.

Mesos


Other big news was the Google Cloud Platform Live conference in SF on Nov 1. The message from #GCPLive was largely about containers… in short, the notion of The datacenter IS the computer going mainstream. To paraphrase one comment during the conf: “Customers get locked into host-based patterns, so they struggle with intertwined systems.” Well said. Definitely looking forward to the new GKE service based on Kubernetes.

Other big news was awaiting in London. Namely, the team behind Weave. Recall that the JEM tutorial had been an alpha test for the IPython + Docker + Rackspace + Nature magazine thing? We learned a truism the hard way, with minutes to go before the event started: Docker does little to resolve crucial issues outside of the containers. Enter Weave, handling difficult matters outside the container, such as networking and crypto. Check their blog for tasty insights, e.g., Automated provisioning of multi-cloud weave network with Terraform. Highly recommended.

Speaking of Docker, I really enjoyed this talk by Adrian Cockcroft @DockerCon: State of the Art in Microservices. Especially slides #8–19, product development process.

Speaking of Microservices, here’s a good overview: The Strengths and Weaknesses of Microservices by Abel Avram on InfoQ.

Ag+Data

Continuing on the Ag+Data front, check out the excellent article GeoTrellis Adapts to Climate Change and Spark about how Climate Change analytics drove Spark adoption at Azavea. They integrated Spark and Accumulo to support fast computation of climate impact metrics for DoE, which should be included in the 0.10 release of GeoTrellis.

NYT ran an interactive analysis/visualization, Flooding Risk From Climate Change, Country by Country, which perhaps helps explain Silicon Valley rumors about Google building ferry ports at corporate campuses along SF Bay.

I’m a big fan of Danielle Nierenberg @FoodTank in Chicago. A recent article, How Vegetables Can Save the World, is brief, accessible, and quite to the point. More of that on FoodTank.
Meanwhile, considering the many challenges ahead in Ag worldwide, I’m curious whether some programmable matter could become useful on farms to leverage data? Sort of an asymptote for IoT.

Upcoming Events

Many interesting conferences and other events are planned for the months ahead. Please do check the http://goo.gl/2YqJZK listings. In particular, mark your calendars for:
O'Reilly studio in Sebastopol, for new "Intro Spark" video

Misc.

I’ll leave you with something fun and something epic.

First, the fun – though it’s quite epic in a way: LumiGeekWe make Arduino shields for LEDs, audio-reactive drivers, and custom solutions for architectural and artistic endeavors. Check their installation at the new Galvanize Cafe in SF, and look about carefully for a subtle case of anamorphosis.

Second, the epic – if you haven’t seen it yet, it’s well worth four gorgeous minutes of video: Wanderers by Erik Wernquist, narrated by Carl Sagan. Money quote @1:45: “Herman Melville in Moby Dick spoke for wanderers in all epochs and meridians…”

That's the update for now. See you in Austin, San Jose, and NYC on the event horizon!

2014-07-27

Newsletter Updates for July 2014

Two aspects about leveraging machine learning are largely under-represented in the lit, especially when it comes to production use cases: feature engineering and the comparative evaluation of multiple modeling approaches. To that point, check out “Streamlining feature engineering: Researchers and startups are building tools that enable feature discovery” by Ben Lorica. The article mentions Spark Beyond, which “finds deep patterns in your data.” I was lucky to get a demo of Spark Beyond earlier this year and talk with the principals – and highly recommend taking a good look at their wares. Between the ongoing advances in deep learning and symbolic regression, a direction seems to be emerging … that perhaps one of the more difficult parts of machine learning workflows, namely the feature engineering aspects, could become more automated.

For another great article, check out Including Men in the Conversation About Women by Scarlett Sieber. Among my biggest peeves about Silicon Valley are the “brogrammer” lopsided demographics, and the gender bias which is quite real and nearly epidemic. Our data science teams have generally been quite mixed, why can’t engineering teams in general leave the 19th century behind, let alone stop being so hostile? Not naming names, but two of the SV firms in which I’ve worked in the past five years are both well known and well poised for harassment lawsuits. Taking a stand against that nonsense as an engineering manager is a great way to catch hell, which I’ve gladly engaged before. Another related pet peeve is where one of the same firms was actively pressuring their engineering interns to quit university degree programs. As a behavior for an engineering manager, I find that highly unethical. Some of those who are engaged in these practices know quite well who I’m talking about.

Spark Summit

The big, BIG news last month was … (wait for it) … Spark Summit. All of the speaker videos have been posted – those are probably the single-best resource for learning about Apache Spark. Of course, the big surprise at the conf was the announcement of Databricks Cloud. If you missed the conf, you can watch Ali Ghodsi’s spectacular demo which kicks in at about the 14:40 time marker.

Spark Summit keynote practice, T-15 hours
One surprise learning from the conf was that one product line from SAP generates more annual revenue than all of the other Big Data vendors (HW, Cloudera, etc.) combined. Other pleasant surprises included: Flambo, a Clojure DSL for Spark; and Thunder, for large-scale neural data analysis, which shows some excellent integration of PySpark, SciPy, scikit-learn, etc.

Our training sessions at Spark Summit set some kind of new records. In particular, check out the advanced material for great lectures there. Those who attended the conf received a free ebook preview for the upcoming Learning Spark: Lightning-Fast Big Data Analytics by Holden Karau, Andy Konwinski, Patrick Wendell, Matei Zaharia; O’Reilly Media (2014).

Also, I got to host the Research track of session talks at Spark Summit, which was a real treat. We had a special #geo break-out session following the Geotrellis talk by Rob Emanuele. We will hopefully be expanding that focus in future confs. There were so many other great talks that it’s hard to pick favorites. Even so, I’ll be studying up about two in particular: Quadratic Programing Solver for Non-negative Matrix Factorization with Spark by Debasish Das, Santanu Das; and Distributed Reinforcement Learning for Electricity Market Bidding with Spark by Vijay Srinivas Agneeswaran, Vishnuteja Nanduri. The latter seems almost ideal for integration with recent work on genetic programming.

Stay tuned for the next Spark Summit, which will be held on NYC in early 2015.

OSCON

I’ve just returned from OSCON 2014. What an excellent conference! Check out the content recently posted online: keynotesphotosspeaker slides.

Of course, this event was carefully timed to overlap with the Oregon Brewer’s Festival. Top two picks: Double Latte. by Sierra Nevada Brewing Co.; and Lorenzini Blood Orange Double IPA by Maui Brewing Company. Many thanks to Erin Rasmussen for suggesting about OBF!

Blood Orange IPA
Back at OSCON… one of my favorite Ignite talks was What Science Fiction Can Teach Us About Building Communities by Dawn Foster. Another favorite, speaking of #geo, was a preso/proposal for Open Aerial Map by Kate Chapman.

During the conf, Andy Orem did a video interview where we discussed perspectives and current projects: Ag+Data, Industrial Internet, sketch algorithms, Apache Spark, etc. Andy was the very first editor I worked with at O’Reilly Media, ten years ago. He’s a much better interviewer than I am an interviewee, so I enjoyed learning much through our work together. Also fun to work again with the amazing video team.
"With great power comes some data, plus wrinkled shirts"
The tutorial for Just Enough Math had 50+ people attending, and we got to evaluate an intermediate stage of a new tutorial software platform. For that, I needed to get a bunch of USB drives from Amazon, but the order/delivery #failed. At the last minute our 10 y.o. daughter and I made an emergency run to Fry’s Electronics (she was eager to observe ground zero for nerdliness) … but the only 4Gb flash drives that they had left in stock were Marvel Universe comix characters. Arriving back home, our 9 y.o. daughter was aghast that adults would be receiving comix figures in a lecture :)

The Data Workflows for Machine Learning talk received lots of great responses – as did earlier versions during meetups in Seattle and SF. It become of the “top-shared” slide decks featured on the SlideShare home page. Perhaps that needs to be turned into a mini-book?

new book kiosk
As my last-o’-the-day book signing was winding down, after almost everyone had left the convention center for “nearby locations of beer taps”, a friend mentioned “Hey, look there’s another pile of books – these look different.” So a few lucky latecomers got signed copies of the galley drafts for our new book Just Enough Math, which probably still won’t be released for months – this rev is quite rough :) Oddly enough, the first person to read it looked up and said, “Where are the other O’Reilly books about math?” Indeed.

Sketchy Things

Speaking of Just Enough Math, we’ve put up a companion site for the video+book+tutorial at http://justenoughmath.com/ to provide additional resources and related links:
  • set up a Python programming envon your laptop
  • code+data files for examples in the video+book
  • “gists” that show expected results for the examples
  • links to external resources that get referenced
  • recommended books and videos for further study
  • monthly newsletter sign-up
The tutorial at OSCON previewed a new chapter recently added about sketch algorithms, following from notes at an excellent Foo Camp session led by Avi Bryant. I will be focussing on Spark Streaming use cases for Strata EU in Barcelona this fall, particularly where approximation techniques (think: examples of monoids in action) can leverage both Spark and Cassandra. If you have examples to share of Spark Streaming production use cases in general, I’m eager to build case studies to publish in Radar. Meanwhile, for a great resource about sketch algorithms, check out the archives of the AK Data Science Summit – Streaming and Sketching from last summer.

Card-Carrying Green

A friend recently brought up the topic of navigating questions about extinction and climate change for preschoolers… I’m getting those too; however, in my experience the questions become much better formulated after an additional 5–6 years or so. As a parent, as a human, it kills me to see all the ginormous FUD spewing from the political lobbies for the coal industry, fracking, Monsanto, GM, etc. How about giving ample air time and consideration for some points from the other side?

First off, I’ve mentioned it before but it bears repeating: The Land Institute is a phenomenally excellent resource for understanding some of the insanity and pure tragedy of contemporary agricultural practices, particularly when it comes to monocultures, annuals, hybrids, let alone unnecessary tillage. To paraphrase Wes Jackson, “The plow share has destroyed more options for future generations than the sword.” On a related note, I’ll also point to an excellent article by Michael Pollan, as a forward to Grass, Soil, Hope: A Journey through Carbon Country by Courtney White. Moreover, check out The Solutions Project. That latter site has more substance than perhaps its web-design polish indicates: it’s about the work by Mark Jacobson, et al., on how to power the planet via renewables now while mitigating hurricane damages, etc. One would think that the reinsurance revenues alone would justify a significant investment. In any case, these three links point to the fact that any emerging “dialog of despair” about global warming, etc., is purely FUD. Much can and will be done.

Phylo, the trading card game
I’m particularly grateful to be associated with O’Reilly Media, which provided OSCON attendees with a nice treat in their schwag bags: Phylo, a trading card game. Its gameplay emphasizes endangered species, climate change, food chains, and other environmental pressures. “Phylo is a project that began as a reaction to the following nugget of information: Kids know more about Pokemon creatures than they do about real creatures. We think there’s something wrong with that. Apparently, so do many others.”

In a related development, check out Nerds Without Borders: “We are looking for all sorts of people to help: Engineers, Scientists, Writers, Artists, Dreamers, Activists, Organizers, Fundraisers, Financiers, etc…” Starting with use of IoT sensors and cell phone networks to protect sea turtle hatchlings. Good stuff.

Looking Ahead

Another fun follow-up from Foo Camp and OSCON: getting to talk with Scott Jenson about his work on The Physical Web at Google. Check out his preso, Why Mobile Apps Must Die. The big idea is a kind of “micro-DNS” for low-cost digital tagging of physical items that can be accessed by mobile devices. No app installs required.

In other news, Trafodion was recently released as open source by HP. The name is based on the Welsh word for “transaction”. If you recall about Tandem Computers and NonStop, this product line has a long history of tech innovations – for highly reliable, highly optimized real-time SQL at scale. My uncle retired from Tandem, and lately I’ve spent time with the Trafodion team and am quite impressed. This release brings an interesting new level of Enterprise robustness to real-time transactions+analysis atop Linux+Hadoop. One to watch.

Another to watch closely is The Distributed Developer Stack Field Guide by Andrew Odewahn, Courtney Nash, Mike Loukides, et al. This is a GitHub-based book from O’Reilly. If you see any points in there that need editing, embellishing, etc., then two words: pull request, for the win.

In terms of upcoming events, registration is now open for Data Day Texas 2015, and I’m really looking forward to that. Will be teaching Spark at Scala by the Bay in SF on Aug 8–9, speaking at #MesosCon in Chicago on Aug 21, followed by another Spark course in Chicago on Aug 25.

Flashbacks

I’ll close with a look back to a 1990 Documentary about Cyberpunk. That provides a good summary of what we up to in the early 1990s with Mondo 2000, bOING-bOING, FringeWare, WiReD, The WELL, Turkey City, etc. Tim’s monologue around 15:30-ff is hilarious – both because of his ever-optimistic “There will be mass democracy in the streets” miss, and how much it contrasts with just about every other major point coming true within 25 years. Warning: gratuitous F242 clips, throughout. Time marker 27:11 shows what I was doing as a vendor at many, many raves… Meanwhile, check out a recent bOING-bOING article Alien Autopsy: William Barker on Schwa, two decades later for some of the more astute counterpoint about what was really going on, then and now.


That's the update for now. See you in Chicago with San Diego on the event horizon!

2012-06-22

hadoop summit 2012: emergence of the confidence economy


moore intro

Geoffrey Moore opened his keynote at Hadoop Summit 2012 and promptly dropped the line: “You will remember this moment years from now.”

After a disappointing set of “sales pitch” keynotes on the first day of the conference (thanks Yahoo! — but you knew that already) many people attending seemed to roll their eyes about yet another keynote talk this morning. Surprise!

I was grateful to hear Geoffrey Moore trash Advertising as an industry at risk. If I may paraphrase: permanently caught between bleeding edge and dinosaurs, yet irreparably dependent on a broken business model. [FWIW, the last three VCs on whom I’ve used that line looked back at me like I was some kind of alien slime-mold.]

By the middle of his talk, Moore put up a slide with a half-dozen bullet points. The slide listed some of the most disruptive technologies on which businesses — Main Street, in his terms — would come to rely in the early 21st century. Those include: collab filters, behavioral targeting, predictive analytics, fraud detection, time series, etc., etc. Outside of the intelligence community and the hedge funds, the significance of these technologies is not well understood yet. Word. Up. Bitches.



Moore’s “Final Thoughts” slide really hit home. He talked about data access patterns (system of record vs. log file usage vs. real-time analytics vs. etc.) and how those access patterns create feedback loops within an organization. Moore claimed this was core DNA for Google, Amazon, etc., which all major businesses must now embrace. Or else. [That's about 95% overlap with a slide I made for (insert recent past employer) during a 2011Q1 pivot. Two pivots later, I left without any particular next gig in mind — clearly needing to get involved with a different business team. Shortly before their CEO got, um, an "opportunity" to find work elsewhere. But I digress.]

an exercise

So here’s a fun exercise for the interested reader: Pull up a 10-year chart for the S&P 500. Add to that CBS. Right.. Add to that Barnes & Noble. Bokay.. Add to that Wal-Mart. Got few bumps, some upturns.. Nothing to write home about.

Now add Google. Now add Amazon. Now add Apple. One might argue that I’m cherry-picking examples; however, one must understand those three in particular to grasp the trajectory of how Data modifies Companies.

Think about it. Imagine rolling the clock back about 13 years, just a few years before that huge financial sea change got going. Think about perceptions at the time of Apple, Amazon, Google. Most of the mainstream buzz that I heard or read in 1999 was largely disparaging about those three. They didn’t make sense to the average joe, and that was a problem. I will contend that what made sense to a handful of computer science grad students, but not to the average joe, was considered a problem for Main Street. A multi-billion dollar existential problem for some, as it turned out.

At the time, it seemed like Apple would never get past the overwhelming popularity of Dell and Microsoft. Amazon didn’t have a way to justify its enormous P/E ratio — and was probably fluff in the long run. Google was considered interesting, but a little strange, with no clear path toward revenue.

Now think about what happened to the music industry, the mobile industry, the … well, I could go on, but Apple disrupted the pants off lots of established players. Entire industries were taken down by one company. Then consider what happened to retail. One word, a verb according to Geoffrey Moore: Amazon. Think about what happened to advertising. Googled, and not in a nice way either. Amazon and Google took off in 1997Q4 and 1998Q1 respectively, with Big Data projects which became enormous cash cows: Amazon’s recommender system (plus cloud infrastructure), and Google’s search+ads (plus cloud infrastructure). Arguably, those two are the reasons we were having a "Hadoop" conference. Apple perhaps seems less in category; however Apple leveraged mountains of consumer data (plus cloud infrastructure) to drive its smartphones, App Store, etc.

Imagine what kinds of conversations which must have been occurring in the board rooms of CBS, Motorola, Barnes & Noble, Wal-Mart, etc., etc. Gone, gone, gone. Three relative underdogs became giants, tipping almost everyone else’s apple carts. (pun intended) At least three firms understood the power of leveraging their data, they understood the urgency of real-time analytics, etc. Their competitors, mostly, did not. Just look at those stock charts.

According to Moore, that was the tip of the iceberg. Most of the Global 1000 is now on notice. Over the next decade we’ll see monumental failures. Winners and losers, as always, but the magnitude of the losers may be unexpected.

central point

Moore’s central point in the keynote — since this was a Hadoop conference — was that the Hadoop tech stack and business ecosystem is maybe a year ahead of the proverbial “crossing the chasm” moment. Ergo his lead line.

Notably, enormous cultural changes of the 1990s and early 2000s have percolated through personal expectations among those coming up in the ranks. That’s happened more notably and with more impact outside the US than within it. He pointed to the “digitization of culture”, where access has become nearly universal, where broadband created emotional dimensions (Facebook, Pinterest, etc.), where mobile makes the experience ubiquitous regardless of socio-economic position.

Meanwhile, the corporate culture of how to “get stuff done” within enterprise has not kept up. There’s no Facebook for enterprise, no YouTube for enterprise, etc. [Well, actually, there are — and they are each headquartered within a bike ride of my home near the Mountain View / Palo Alto border — but you haven’t heard about them. Yet.]

Meanwhile, Facebook-esque consumer Internet companies of the world are too caught up in their own weirdly distorted realities to solve the larger business problems. Business problems where the solutions will inevitably derive from the social networks’ innovations. Oops.

In Moore’s vaulted opinion, those conditions won’t hold much longer. There will be winners. There will be losers. Big ones.

Meanwhile, for people of my ilk, Moore smiled and predicted: “This should provide at least a decade of entertainment for everyone present.” Fundamental business reasons are simple: enormous change ahead but precious few who are trained and experienced to navigate it.

key take-aways

My first key take-away is based on the observation last year that Enterprise giants bumbled into Hadoop Summit 2011 in a huge and awkward way. Oddly, the logo is an elephant, #justsayin

In contrast, this year was really smooth, completely professional, far too expensive … but almost all about data infrastructure in a world where nobody want to utter the word “Oracle”.

Mind you that neither of the two main “enterprise” keynote speakers from last year still have their same jobs. #justsayin

Also, notable Hadoop practitioners were noticeably absent. In fact, most of the cast and crew of Strata seemed to be missing. A particularly popular social network has been burning the midnight oil to make Hadoop perform backflips — they like gave a couple talks and seemed to vanish.

Let me put this in other words: several hundred million dollars have been invested by VCs (and angels) to recreate an industry in the image of Redhat and Yahoo!

Wow, did anybody think that would be a particularly good idea? No, but it’s the herd mentality in practice. Even after the 5th beer I’d still recognize that strategy as not particularly wise. Feels like when you talk with an ex-convict, and they drop a line “Yeah, I made some poor choices long ago…”

My hunch is those data infrastructure plays are mostly tax write-offs (for the “early adopter” part of Geoffrey’s famous curve) at this point.

Moore underscored how real payouts come when key verticals catch fire — with serious domain expertise leveraged. LinkedIn perhaps got close, but now it almost feels like a spamming broadcast system for HR and BD departments. We’ll see “Big Data” killer apps which mean something to lots of people. Beyond the GOOG+AMZN+AAPL tip o’ the iceberg. They will come from people who have sophisticated backgrounds in Stats + ORSA + distributed systems + functional programming + DevOps, people who can also communicate well with actual business leaders. Not those employed by some halfwit B-school grad who’s posturing as the next Steve Jobs, when in reality he drinks bad beer at a lame, faux-hipster sports bar while watching cable televison. Or something. Dude, hop on your fixed-gear bike and standstill/peddle your sleeve tattoos out of here.

Translated: the proverbial ignite moment, that spark of innovation, is not going to come from the likes of a Cloudera or a Platfora or a (banal noun)-(o|e)ra… But it’s going to come, probably not many moons away. It will be in apps.

the sound of disruption

Thirty years ago, I went into a field called “math science”, i.e. how to build predictive analytics as software apps. Stanford — the Statistics department chairman, Bradley Efron, in particular — had put together an interdisciplinary degree which combined math, statistics, operations research, programming, engineering, etc. At the time, most of my peers in the program went on to become insurance actuaries. I went instead to do graduate work in machine learning and distributed systems.

For nearly two decades, most employers could care less about any quantitative background. They wanted C++ software engineers working all day on APIs from Sun or Microsoft or Oracle. Or they wanted managers. Then, in about 2000, came the sea change.

Right about the same time as ticker symbols for Apple and Google and Amazon were strolling up to their respective launchpads, some people began to look at my resume and ask a different line of questions.

I’ll always remember the first: a microchip vendor — one which makes electronics for several products you’ve purchased — was getting squeezed by Intel and their silicon compiler vendor. Critical features were being deprecated, specifically to put this second-tier player out of business. The company was on notice. They had to find a proverbial needle in a haystack: out of tens of thousands of circuit designs, they had to identify the 1% which would no longer be licensed — then redesign those. Quickly.

An internal team at the company had tried, but given up. Too much data for their techniques, it would’ve taken years to resolve. The company hired an electronics consulting firm in Austin, and engineers went to work, but gave up as well. Too much data, not enough signal. I got called in, as a “Whatever, just see if you can get anywhere” last-ditch effort. About 20 lines of Perl and one relatively simple equation later, I dumped my results into a scatterplot.

One of the lead circuit designers picked up my plot off the laser printer and began laughing. Loudly. The whole office heard him.

His manager grew annoyed: “What?! Why are you laughing?!”

Engineer: “He found it.”

When I turned in my invoice, the manager glared. “Look,” he said in a growl, “Just go somewhere for about three weeks. Bill us the whole time. Then come back and turn that in.”

My brows furrowed, this was a high-dollar rate for 2000.

“If you don’t pad that damn invoice…” he paused, “You’ll make both us and our customer look like complete fools. Piss a lot of people off.”

That’s the sound of Disruption.

More than a decade later, the summary graf of my resume reads like bullet points from Geoffrey Moore’s slide. Collab filters, anti-fraud classifiers, predictive analytics, etc. Even in the past few years, when HR people have read that resume, several looked up with a frown, said they thought that kind of work was better suited for business analysts — yadda, yadda, yadda, keep following the herd: you put the “botch” in “beotch”.

At a time when lots of business (start-ups as well as enterprise) are starving because they cannot hire Data Scientists, I’ve been busy building teams. Teams which delivered $MM results. I’ve hired about thirty people onto Data Science teams within the past few years — at a time when many start-ups would feel lucky to hire one. #justsayin

mal*wart

I read one of the most imbecilic essays recently from Forbes/Quora: “What Would Be The Global Impact If Wal-Mart Abruptly Shut Down?” Essentially, a hagiography stating that Wal-Mart is too big to fail, that the consequences on the US economy, the global economy, would be catastrophic. Translated: may require an enormous bailout, soon.

[In case you hadn’t guessed, I just threw up a little bit in my mouth.]

What. A. Fucking. Moron. The reality is that Wal-Mart hasn’t been doing so well over the past decade. Not if you peel back enough layers of PR. Not since they tried to bamboozle the LA city council. And failed. Moreover, folks at Amazon could really care less which Senators or SecState/former-first-lady the execs in Arkansas have in pocket. Bezos has positioned to take over 150% of Wal-Mart’s business the picosecond after Bentonville implodes. Sears and Target have reinvented themselves specifically for that very instant. So long, good riddance. Remember the point about the Global 1000 on notice? About the importance of business fundamentals?

sears, a.k.a. that web site which kinda looks like amazon

A third keynote talk that day was by the Sears CTO, Phillip Shelley. I had packed up my laptop and backpack, and was getting ready to walk out of the auditorium. After his first few sentences, I put my stuff back down and started taking notes.

Dr. Shelley mentioned how Sears started as a mail-order business a century ago, though more recently got completely kicked by another “catalog” called Amazon. Now they are leveraging Hadoop + R + Linux/Xen private cloud (srsly, this is from the Sears CTO?!?) to reinvent their business with 100x more detail on regional pricing models. Literally calculating personal pricing discounts for individuals, multiple times per day, specifically for mobile.

Sears: core algorithm moved from [6000 lines of COBOL on mainframe with 3.5 hr batch window] to [50 lines of Hadoop app on Linux with 8 min batch window], while reducing TCO for enterprise IT by two orders of magnitude. So much success, that they’ve spun it out as a new business line called MetaScale.

Brilliant strategery by Sears. Some seriously high powered Data Science talent walked out of that auditorium musing how they wished their VP Engineering was half as progressive as Sears. Srsly? Um, that’s called a PR coup. [Literally at the same moment as Wal*Mart had HR droids spamming the audience with whispers and rumors of lucrative salaries. Gak.]

emergence of the confidence economy

What’s the deal? It’s about confidence. Those giants in the Global 1000 which Geoffrey Moore says are on notice? They got that way by believing that business is largely about who barks the loudest, barks the longest, and cuts the most deals under the table. The proverbial alpha male in a wolf pack.

Wal-Mart would be a prime example, in my opinion. Their business is predicated on fundamentals which simply do not hold. Misplaced confidence. Thanks to people like Hillary Clinton, Wal-Mart has gained much influence on the House and Senate floors and the halls of the State Department and the UN assembly. In other words, so long as we manage to keep fuel costs artificially low, Wal-Mart’s market valuation will keep growing. So long as we believe that bullying vendors, conducting intelligence operations against the rest of your ecosystem, etc. — that these kinds of practices are ethical and sound in the long-run, then Wal-Mart will keep growing. Bullshit. Go look at that stock chart again. Wal-Mart is about tall white guys in dark suits, acting like complete pricks, destroying and plundering anything they can get their grimy paws on. Richard Gere in Pretty Woman, before he gets Julia Roberts. And not much more than that. On notice.

Moore is pointing out, in my opinion, that the issue at hand is about uncertainty. The point of establishing a corporate charter was always to externalize risk and perpetuate wealth for shareholders. That was true four centuries ago, when the first transnational was established, and has been true ever since. The modus of that mechanism is a process called sublation. The train wreck for sublation is uncertainty. In an environment where uncertainty holds sway, having real-time analytics from petabytes of customer data wins out over having a Senator in pocket. Any day of the week. The antidote for uncertainty is confidence. While there had been a regime of an “Attention Economy” extant for the past two decades or so, we’re now entering a new regime of the “Confidence Economy”.

Here’s the deal: people like me like those of us in Moore's lecture have been the “secret sauce” fueling the rise of Amazon, Google, Apple, etc. We use techniques which are mostly not well understood outside of Langley and the hedge funds. The tools of contemporary corporate assassins. Guys in suits who act like pricks in lieu of practicing business fundamentals — those guys are our targets. The modus is Disruption. If you have an MBA or a CxO title and not much else to back it up, I put food on my family’s table by being a sniper paid to hunt you. Lots of *great* food. And some of the best wines available. I shake the tension out of my hands, correct for wind and distance, draw a bead, take a deep breath, squeeze the trigger. Kill shot.

The challenges faced by Data Scientists are daunting. On one hand, most mathematicians lack enough solid engineering to create killer apps. Conversely, most engineers lack enough math to make any headway on the business data. Most business analysts lack enough of either the math or the engineering to be worth hiring. Data Scientists provide all three areas of expertise: the engineering and the math and the business insights to contend with mountainous torrents of data, and move the needle. On the other hand, Data Scientists must also speak truth to power. In any given business, there will be winners and losers. Executives, people accustomed to their own power, taken down. Meanwhile, we Data Scientists come prancing into a business, we do our magic, and consequently we point out which executives are bullshit and must be “executed”.

The reason why I’ve built Data teams at a time when others are starving is simple: confidence. Sure, I’ve logged three decades of machine learning, statistical modeling, data management, distributed computing, etc. When I talk with a grad student about their work, I can tell them in 25 words or less what they need to do on their first day at work to become regarded as an great asset to the team. They already know the techniques, but crave confidence. Into the trenches, fresh-out, having to speak truth to power. They’ll be placed into some faltering business unit, run some detailed analysis, and point out that the VP who’s been arguing loudly was completely wrong for the last N years and his/her ego cost the company several $MM. You can bet that those execs will return fire. However, a person like me is confident that we can get a kill shot. I show new folks how to draw a bead and squeeze the trigger. Been doing it for a long while, and will be doing for a long while more.

Snipers have an eerily pragmatic sense of confidence. And, by the way, that’s a peculiarly difficult job. Praise goes out to the men and women who serve their countries in uniform — when the cause is just. [FWIW, before tackling the challenges of data+science, I wore a military uniform and carried a rifle. Sniper training has become invaluable.]

my opinions

#1: Enterprise suffers because so many people in the corporate leadership ranks (or rather, amongst those clawing and scrambling to make their way into the corporate leadership ranks) consider themselves to be a different caste — if not a different species all together — from the rest of us who do not have a salaried position with a transnational. In a “Let them eat cake” world fraught with trillion-dollar bailouts, that’s not a particularly good way to future-proof. Moreover, this is why VCs are vital… to demolish that kind of hubris via constructive Disruption. #justsayin Word. Up. Bitches.

#2: Enterprise tooling, which is now mostly dependent on JVM-based apps, suffers because it has embraced “Convention over Configuration” … CoC has its place. I can imagine that it’s an excellent idea for heart surgeons to have a standard toolset, with scalpels in the exact same positions, etc. CoC is not a particularly good way to manage complexity and uncertainty, because it simply displaces major problems into the build system. Ultimately, it fails too much and impedes spin-up. Which, I believe, represents an enormous, ticking time bomb in Enterprise. Here’s a challenge: Show me a metric for the median period it takes in your business for a newly hired engineer to push code changes into production use which is adopted by at least 80% of your customer base. Now show me a metric for the media period it takes in your business between the point where a product manager identifies a needed feature and a newly hired engineer is ready for spin-up. From those, I’ll make a prediction based on that metric for how well your business will survive the “on notice” condition which Geoffrey Moore described. Better clues for navigating complexity and uncertainty can be found in the works of Ilya Prigogine or Stephen Wolfram. To wit, functional programming is more likely to address complexity, real complexity, and also more likely to attract top talent. CoC, not so much. Perhaps your enterprise business addresses Main Street instead of Early Adopters... Recall that Google and Amazon and Apple crossed the chasm by recruiting armies from grad students — at a time when most other people erred on the side of average joes. Remember the point about real-time analytics? It counts for training your people, then retraining, and retraining, constantly — to grapple with uncertainty. Kill shot. Global 1000.

#3: MapReduce will be unrecognizable within three years. Hadoop Summit will become something quite different after Hadoop bifurcates and gets sublated into Something Else. For example, it would be not difficult to use the Simple Workflow Service from Amazon AWS to implement MapReduce using the core part of Cascading… a different kind of MapReduce, which is not constrained by JVMs… which could scale much more gracefully and robustly… which could out-perform Google infrastructure and avoid attempting to re-create the industry in the image of Yahoo! At which point, one could deploy functional programming blocks at enterprise scale, without having to rely on the morass of enterprise build tools. Hmmm… may need to get a term sheet for that one.

Geoffrey Moore, we may have a few answers for your questions.