Showing posts with label stanford. Show all posts
Showing posts with label stanford. Show all posts

2015-03-07

SV Synopsis: Fundamentalism in Technology

I am grateful for perspectives gained because our family lives in Silicon Valley. Many options here to work at novel ventures, and on fascinating projects… Opportunities to drop by Stanford or Berkeley for some remarkable guest lecture by a visiting expert… The wonders of an almost perpetual Maker Faire as one walks through the neighborhood on any given evening… Tech camps that our daughters can attend locally, as they wish… And, generally speaking, the lack of any real need to engage in ridiculous commutes

As an open source evangelist and as an investor, I've felt grateful to learn from a veritable parade of interesting projects. However, I am troubled by the incidence of a particular problem. Far too often one runs headlong into what I could characterize as a close approximation of cocaine-fueled misogynistic narcissism. The condition is subtle, but systemic here. Even recently, I have witnessed this up close – along with the regrettably pervasive and predictable non-reactions to it. Increasingly, zero tolerance appears to be the only effective response. Or perhaps the tech industry percolates out elsewhere, far from SF and its inertia?

Without mentioning names, two well-known billionaire-club investors in Silicon Valley personify this character sketch. Evidence of panspermia ad absurdum festers in the "cultures" that they promote. Personal jihads seemingly to self-perpetuate their fundamentalist ideals.

A nagging question lingers… Why work alongside an ilk of people with whom I would never encourage my daughters to mingle? Granted, I believe quite strongly in the need to talk with just about everyone, to keep dialogue open, to reject the notion of "enemy". Even so, there are absolutes. Practical realities of livelihood aside, as a parent what kind of examples do my professional actions and affiliations set?

In addition, a question that investors ask over and over when considering whether to fund a new company is "Will the team scale?" Any measure of the toxins described above almost guarantees that the answer will in practice be "No."


That represents a dirty little secret. There is an amazing level of demand for tech talent. It's not exactly because these companies are raging commercial successes; most early-stage ventures by definition are not. It's because few people who are capable of making good judgements are willing to compromise their futures to work for ineffective caricatures. Many start-ups encounter difficulties in scaling their team. Or – more likely over time – they encounter high attrition rates.

While I have in the past focused for several years on the same project, lately I don't stay long in most early-stage firms, generally moving on after an organization demonstrates its nature. To paraphrase Lady Grantham from Downton Abbey, there is a point at which malice ceases to be amusing. On the one hand, that's a terrible way to leverage stock option packages. On the other hand, arguably I have pursued a portfolio career strategy. That approach has helped me build an amazing network. Long-term benefits of my network have far surpassed the potential upside of my aggregate stock options. Therein dwells an important lesson about Silicon Valley.

2014-09-30

Newsletter Updates for September 2014

Highly recommended, Oct 2: an O’Reilly Media webcast Spark 1.1 and Beyond by Patrick Wendell and Ben Lorica. Two people who have much to share about where Apache Spark is heading.

My favorite conference in a long while was the Spark Tutorial hosted by Prof. Reza Zadeh @ Stanford ICME – home of world-leading innovation for machine learning at scale. The tutorial featured lectures on Spark Streaming, MLlib, GraphX, etc., from lead committers. Great to be working at Stanford again (if only for a few days this summer) and wonderful to meet many people who participated. Here’s an excellent set of notes. For Stanford affiliates, Prof. Zadeh has an upcoming course CME 323: Distributed Algorithms and Optimization with related content explored in much more detail.

We will hold another Spark Tutorial at UMD in College Park, Maryland on Oct 20–22, hosted by Prof. Jimmy Lin. That event sold out quickly, as did the one at Stanford – so we’ll do more! More about that in a bit.

The Quad @ Stanford University
Another great conference this summer was the inaugural MesosCon 2014 in Chicago last month. Twitter kindly recorded all the sessions. In particular, Ben Hindman’s keynote hints toward cross-datacenter features on the horizon. My talk was about Spark on Mesos, and a related blog post shows a few simple steps to launch a Spark cluster on Mesosphere’s free-tier service atop Google Cloud Platform.

Mesosphere partnered with Google’s Omega team for a killer demo involving Kubernetes and Mesos, showing cluster failover/migration across datacenters in CA and NY. Sounds simple, but the implications are vast. The other killer demo, from eBay, featured YARN on Mesos – with ultimately no code mods required, just an additional JAR file plus some config settings. Check out related slides and video. Ginormous implications for that one, thanks eBay!

Sparky-the-Bear sez: ignite your data

Big news for me this summer was joining Databricks as Director of Community Evangelism. New business cards. Lotsa new tshirts. I’m thrilled to become part of this renowned team, delighted to be out in the field amidst the exponential growth of Spark production use cases.

KDnuggets ran a story recently about our Spark news… and there’s a lot. To quote the Gartner report Hype Cycle for Advanced Analytics and Data Science 2014: “Databricks is providing certification, training and evangelism that mirror the early Hadoop model.” Of course AMPLab + Databricks have been running Spark training sessions for years. I’ve joined to lead this program, and our team is busy delivering:
Databricks and O’Reilly Media partnered to launch Developer Certification for Apache Spark http://oreilly.com/go/sparkcert – a brand spanking new program that leverages the amazing Spark experts @ Databricks + the incomparable editorial team @ O’Reilly Media:

val results = sc.parallelize(world_class).map(x => exp(log(x) * 2))
results.sum()

So my second O’Reilly book turned out to be a video + Docker image, while the third became a cert exam :) This formal exam takes < 90 minutes: expect multiple-choice questions based on small blocks of code in Python, Java, Scala. Questions test for a range of developer knowledge across Spark Core plus Spark SQL, Streaming, MLlib, GraphX, and typical use cases. We’re establishing the industry standard for measuring and validating technical expertise in Spark.

How to prep for this exam? Don’t worry, it doesn’t require extensive Scala knowledge; however, some familiarity with Scala code examples shown in the Spark docs would help lots. Mostly, we’re testing to see if you understand the Spark execution model, RDDs, how to leverage functional programming to get the most out of your cluster, i.e., avoid common bottlenecks, refute some of the, ahem, FUD that’s been circulating about MapReduce vs. Spark. You are probably good to go if you:
Alternatively, we’re looking for volunteers. The certificate exam will preview on Oct 16 at Strata NY and we need volunteers to evaluate the exam. You’ll get deep discounts on the Spark developer certificate. Plus, it’s an excellent way to score ginormous brownie points with both Databricks and O’Reilly Media, along with conf coupons, outstanding nerd cred, etc. Become an essential part of the Spark developer community building the next-generation of Big Data apps. Let me know. I’ve heard that T. O’Reilly and I. Stoica have authorized us to buy NY gourmet pizza + top-shelf beers for all volunteers (at least let’s start the rumor).

Meanwhile, stay up to date with the latest advances and training in Spark, and help prep for the certification exam. Workshop materials are authored by Databricks, and we’ve trained and certified these instructors. Upcoming training for Spark will be held in SF, DC, London, Paris, Barcelona, Stockholm, and Dublin:
I look forward to the EU trip, but I regret not arriving in time for Scala.IO – amazing talks lined up this year. Also looking forward to Big Data TechCon, and in particular I recommended The Hitchhiker’s Guide to Machine Learning with Python and @ApacheSpark by Krishna Sankar.

BTW, keep your eyes peeled for more material (courses, talks, videos, webcasts, etc.) about architectural design patterns that leverage Spark together with other popular frameworks, such as Cassandra and Kafka. Our team has been working closely with DataStax and others to bring you solutions that go far, far Beyond Hadoop. For those who weren’t watching closely: an emerging tech stack that integrates Spark, Cassandra, Kafka, ElasticSearch, etc., recently pulled in a 1/4 billion in VC financing.

Just Enough Math

The Just Enough Math material is progressing well… Similar to OSCON, we’ll have a tutorial at Strata NY on Wed, Oct 15 1:30pm, expecting +100 people this time. There’s also a public Docker image now, plus more work with O’Reilly on this project. We needed more Mesos + Docker foo to make progress on that infrastructure.

Hopefully, we’ll have an upcoming series of lectures too!

3D Printer Room @ Singularity University

The return of the fellowships

It was an honor to present at Singularity University this summer, along with a workshop at Insight Data Engineering Fellows Program. Looking forward to visiting Zipfian Academy soon too.
We have bunches and gobs o’ regional confs and meetups scheduled:
Also mark your calendars for:

Ag+Data

Continuing on the prior theme of Ag+Data, James Hamilton (Amazon) wrote an intriguing blog post recently, Data Center Cooling Done Differently about a new kind of collocation: datacenters and desalinization. Desalinization at scale seems inevitable here in California – perhaps taking a cue from successes in Australia, etc. FWIW, I prepared a VC pitch for a related venture in 2008, but pulled back after initial feedback. Remember: always go with your gut!

I thoroughly enjoyed this gem about “Organic Ready” non-GMO seeds… Here's to gametophytic incompatibility in large doses. Also check Water’s Edge for an interesting special report on rising sea levels. Big Data comes in handy for contending with these crises related to global warming issues. Three items to check out from low Earth orbit: The SatelliteSpaceknowOmniEarth. Just in case we fry the biosphere before we can get a semi-permanent backup archived on Luna or Mars… one dreads the thought, but artificial photosynthesis is becoming more of a reality. I say “dread” because that idea recalls a vision of Trantor or perhaps Silent Running.

While we’re talking about remote sensing, I should also mention a follow-up study on the data point about GE 12 exabytes/day from turbine sensors on commercial flights: 2000x faster detection of rare critical failure modes. Here's to those early successes turning into a trendline for IoT.

Misc.

A few pointers to notable work by friends and family: Film Theory and Chatbots by Robby Garner; Don Webb: Writing the Science Fiction Novel @ UCLA Extension; Eisoptrophobia by Akira Rabelais; AlaVoidDistribution by William Barker. 

Then I’ll leave you with something haunting and epic: NASA Space Sounds.


That's the update for now. See you in NY, DC, EU on the event horizon!

2014-05-26

Newsletter Updates for May 2014

Been quite an interesting past month or so: DC, Austin, SF, Ann Arbor, Atlanta, Seattle… with hopefully much learned from those travels, plus many excellent events and introductions.

Meanwhile, I learned much from this gem, Therbligs for data science: A nuts and bolts framework for accelerating data work, by Abe Gong. Looking forward to seeing more about Therbligs from Abe. Definitely tune in to Welcome to Intelligence Matters, a new series by O’Reilly exploring current issues in AI, with Beau Cronin as lead correspondent. Another recommended gem is Genomics Crash Course for Data Engineers by Allen Day – that's at the intersection of Genomics and Big Data, for which I have seen an uptick recently.

Just Enough Math

Allen and I have been working to complete our new O’Reilly book, Just Enough Math. The video is in post-production now, and the book is half through second drafts – we are closing in! Some of that material will be previewed in the upcoming workshop Machine Learning for Managers:
O’Reilly will host a free one-hour webcast, Computational Thinking, Just Enough Math on Wed, Jun 4, 10:00am–11:00am (Pacific). Please join me there. The webcast will help publicize a tutorial based on Just Enough Math at OSCON in Portland on Sun, 20 Jul, 9:00am-noon. As a special offer, use the code PACOID to get a 20% discount on OSCON registration. Our tutorial will preview a very new thing at O’Reilly: converting book+video content into interactive tutorials using Docker + IPython Notebook + Vagrant + Git for a cloud-based next-generation content platform.

Speaking of Docker, one of the more interesting start-ups that I have run across recently is Resin, using Docker and Git to containerize+push apps on IoT devices running embedded Linux. Brilliant work.

UCB Initiation Ritual: cousins circa 1968, near Atascadero

In other news, I am thrilled to announce a partnership with Databricks, where I’ve been working to help develop an instructional program that introduces Apache Spark. As you can see in the photo above, the ceremonial ritual for teaming up with UC Berkeley is a bit arduous, but well worth it. Yes, you heard correctly … a Stanford alum saying “Go Bears!”

Our first course in the series is Databricks Hands-on Intro to Apache Spark, an introduction for developers working in Python, Java, and Scala. We have several of these workshops scheduled:
Spark is approaching the 1.0 release at Apache, with new support for SQL. Overall, one of the best presentations that I’ve seen recently about it was Spark at Twitter by Sriram Krishnan, Engineering Manager for Data Platform at Twitter.

The agenda was posted recently for Spark Summit 2014, in SF on 30 Jun - 1 Jul. As another special offer, use the code Paco2014 to get a 15% discount on Spark Summit registration. Highly recommended, and I hope to see you there.

Mesos Updates

Speaking of BDAS and the Berkeley Stack… there have been lots of developments in the Apache Mesos world. One of the best talks ever about Mesos was Improving Resource Efficiency with Apache Mesos by Christina Delimitrou, a case study about Quasar usage at Twitter. Also check out Mesos Elastically Scalable Operations, Simplified by Niklas Nielsen and Adam Bordelon, presented recently at ApacheCon 2014.

The other big news is that #MesosCon, the first Mesos conference, will be held in Chicago on Aug 21. Definitely see you there! Companies interested in sponsoring the conference – please inquire.
I’ve create a new workshop called Cluster Compute App Integrations about building end-to-end apps for Big Data. The workshop leverages Mesos based on the https://elastic.mesosphere.io/ service in the cloud, along with Spark, KNIME, etc. Hint: this involves teams competing, and it is turning out to be quite a popular course. We have upcoming dates lined up:

Agriculture + Data

Did you know that agriculture provides a livelihood for 40% of the world’s population? Or that agriculture consumes 70% of the world’s freshwater in aggregate? That figure is expected to reach 89% by 2050. Or have you heard that Havana grows 75% of its own food based on urban agriculture?
Last month I wrote an O’Reilly Strata article, Ag+Data, about those topics and more. The article introduces a whitepaper, Agriculture + Data: Outlook 2Q14, that we recently at The Data Guild to explore these issues in greater depth. Many thanks to Bill Worzel, Brad Martin, and others who helped on that!

Evolutionary Algorithms

Recently I gave a keynote talk at the Genetic Programming in Theory and Practice conference, which hosted each year at U Michigan by The Center for the Study of Complex Systems. They are the experts in GP; I was merely there to add a few perspectives about machine learning and Big Data. What a wonderful conference. Got to speak at length with Lee Spector at UMass Amherst and Hampshire College. Lee and his grad students have been working with a Clojure-based language called Push, in which evolutionary programs are expressed.

What kinds of optimization problems respond to evolutionary pressure? Definitely not the kinds that one typically finds solved by machine learning. That is where GP approaches come in. In general, there was a lot of discussion about symbolic regression as a general rubric, also some exceptionally interesting work on use of Pareto optimal fronts for model archives (which I’ll be added to my ML bag o’ tricks). In particular, great work from Theresa Kotanchek and Mark Kotanchek at Evolved Analytics. Their software effectively leverages Pareto optimality to select exemplars when models diverge, which I find to be a fascinating alternative to what other disciplines might attempt to resolve through sample. Brilliant work.

Also got to talk with Bill Tozier, author of Answer Factories: The Engineering of Useful Surprises, and viewed some astounding work in HeuristicLab, an interactive framework from HEAL. Think: evolutionary IDE. Another excellent tip was to check out Modeling global temperature changes with genetic programming by Karolina Stanislawska, Krzysztof Krawiec, Zbigniew Kundzewicz.

✽ ✽ ✽

Didn’t get to mention yet about Atlanta, but I really appreciated meeting many wonderful folks there. You’ll be hearing more about upcoming Atlanta plans soon! Also, there are workshops and meetup talks planned now for: NYC, SV/SF, Austin, Chicago. Next up after my current week in Seattle comes Hadoop Summit, on 3–5 Jun in San Jose. Hope to see you there!

-alaVoid Distribution

Misc. Inspiration

In closing… Those who have known me for, well, for the past 20-odd years or so will be familiar with the following: a 21st century artist named William Barker, formerly acclaimed of Schwa Corporation has a new endeavor called -AlaVoid Distribution. Definitely check out his new shop on Etsy.

2012-06-22

hadoop summit 2012: emergence of the confidence economy


moore intro

Geoffrey Moore opened his keynote at Hadoop Summit 2012 and promptly dropped the line: “You will remember this moment years from now.”

After a disappointing set of “sales pitch” keynotes on the first day of the conference (thanks Yahoo! — but you knew that already) many people attending seemed to roll their eyes about yet another keynote talk this morning. Surprise!

I was grateful to hear Geoffrey Moore trash Advertising as an industry at risk. If I may paraphrase: permanently caught between bleeding edge and dinosaurs, yet irreparably dependent on a broken business model. [FWIW, the last three VCs on whom I’ve used that line looked back at me like I was some kind of alien slime-mold.]

By the middle of his talk, Moore put up a slide with a half-dozen bullet points. The slide listed some of the most disruptive technologies on which businesses — Main Street, in his terms — would come to rely in the early 21st century. Those include: collab filters, behavioral targeting, predictive analytics, fraud detection, time series, etc., etc. Outside of the intelligence community and the hedge funds, the significance of these technologies is not well understood yet. Word. Up. Bitches.



Moore’s “Final Thoughts” slide really hit home. He talked about data access patterns (system of record vs. log file usage vs. real-time analytics vs. etc.) and how those access patterns create feedback loops within an organization. Moore claimed this was core DNA for Google, Amazon, etc., which all major businesses must now embrace. Or else. [That's about 95% overlap with a slide I made for (insert recent past employer) during a 2011Q1 pivot. Two pivots later, I left without any particular next gig in mind — clearly needing to get involved with a different business team. Shortly before their CEO got, um, an "opportunity" to find work elsewhere. But I digress.]

an exercise

So here’s a fun exercise for the interested reader: Pull up a 10-year chart for the S&P 500. Add to that CBS. Right.. Add to that Barnes & Noble. Bokay.. Add to that Wal-Mart. Got few bumps, some upturns.. Nothing to write home about.

Now add Google. Now add Amazon. Now add Apple. One might argue that I’m cherry-picking examples; however, one must understand those three in particular to grasp the trajectory of how Data modifies Companies.

Think about it. Imagine rolling the clock back about 13 years, just a few years before that huge financial sea change got going. Think about perceptions at the time of Apple, Amazon, Google. Most of the mainstream buzz that I heard or read in 1999 was largely disparaging about those three. They didn’t make sense to the average joe, and that was a problem. I will contend that what made sense to a handful of computer science grad students, but not to the average joe, was considered a problem for Main Street. A multi-billion dollar existential problem for some, as it turned out.

At the time, it seemed like Apple would never get past the overwhelming popularity of Dell and Microsoft. Amazon didn’t have a way to justify its enormous P/E ratio — and was probably fluff in the long run. Google was considered interesting, but a little strange, with no clear path toward revenue.

Now think about what happened to the music industry, the mobile industry, the … well, I could go on, but Apple disrupted the pants off lots of established players. Entire industries were taken down by one company. Then consider what happened to retail. One word, a verb according to Geoffrey Moore: Amazon. Think about what happened to advertising. Googled, and not in a nice way either. Amazon and Google took off in 1997Q4 and 1998Q1 respectively, with Big Data projects which became enormous cash cows: Amazon’s recommender system (plus cloud infrastructure), and Google’s search+ads (plus cloud infrastructure). Arguably, those two are the reasons we were having a "Hadoop" conference. Apple perhaps seems less in category; however Apple leveraged mountains of consumer data (plus cloud infrastructure) to drive its smartphones, App Store, etc.

Imagine what kinds of conversations which must have been occurring in the board rooms of CBS, Motorola, Barnes & Noble, Wal-Mart, etc., etc. Gone, gone, gone. Three relative underdogs became giants, tipping almost everyone else’s apple carts. (pun intended) At least three firms understood the power of leveraging their data, they understood the urgency of real-time analytics, etc. Their competitors, mostly, did not. Just look at those stock charts.

According to Moore, that was the tip of the iceberg. Most of the Global 1000 is now on notice. Over the next decade we’ll see monumental failures. Winners and losers, as always, but the magnitude of the losers may be unexpected.

central point

Moore’s central point in the keynote — since this was a Hadoop conference — was that the Hadoop tech stack and business ecosystem is maybe a year ahead of the proverbial “crossing the chasm” moment. Ergo his lead line.

Notably, enormous cultural changes of the 1990s and early 2000s have percolated through personal expectations among those coming up in the ranks. That’s happened more notably and with more impact outside the US than within it. He pointed to the “digitization of culture”, where access has become nearly universal, where broadband created emotional dimensions (Facebook, Pinterest, etc.), where mobile makes the experience ubiquitous regardless of socio-economic position.

Meanwhile, the corporate culture of how to “get stuff done” within enterprise has not kept up. There’s no Facebook for enterprise, no YouTube for enterprise, etc. [Well, actually, there are — and they are each headquartered within a bike ride of my home near the Mountain View / Palo Alto border — but you haven’t heard about them. Yet.]

Meanwhile, Facebook-esque consumer Internet companies of the world are too caught up in their own weirdly distorted realities to solve the larger business problems. Business problems where the solutions will inevitably derive from the social networks’ innovations. Oops.

In Moore’s vaulted opinion, those conditions won’t hold much longer. There will be winners. There will be losers. Big ones.

Meanwhile, for people of my ilk, Moore smiled and predicted: “This should provide at least a decade of entertainment for everyone present.” Fundamental business reasons are simple: enormous change ahead but precious few who are trained and experienced to navigate it.

key take-aways

My first key take-away is based on the observation last year that Enterprise giants bumbled into Hadoop Summit 2011 in a huge and awkward way. Oddly, the logo is an elephant, #justsayin

In contrast, this year was really smooth, completely professional, far too expensive … but almost all about data infrastructure in a world where nobody want to utter the word “Oracle”.

Mind you that neither of the two main “enterprise” keynote speakers from last year still have their same jobs. #justsayin

Also, notable Hadoop practitioners were noticeably absent. In fact, most of the cast and crew of Strata seemed to be missing. A particularly popular social network has been burning the midnight oil to make Hadoop perform backflips — they like gave a couple talks and seemed to vanish.

Let me put this in other words: several hundred million dollars have been invested by VCs (and angels) to recreate an industry in the image of Redhat and Yahoo!

Wow, did anybody think that would be a particularly good idea? No, but it’s the herd mentality in practice. Even after the 5th beer I’d still recognize that strategy as not particularly wise. Feels like when you talk with an ex-convict, and they drop a line “Yeah, I made some poor choices long ago…”

My hunch is those data infrastructure plays are mostly tax write-offs (for the “early adopter” part of Geoffrey’s famous curve) at this point.

Moore underscored how real payouts come when key verticals catch fire — with serious domain expertise leveraged. LinkedIn perhaps got close, but now it almost feels like a spamming broadcast system for HR and BD departments. We’ll see “Big Data” killer apps which mean something to lots of people. Beyond the GOOG+AMZN+AAPL tip o’ the iceberg. They will come from people who have sophisticated backgrounds in Stats + ORSA + distributed systems + functional programming + DevOps, people who can also communicate well with actual business leaders. Not those employed by some halfwit B-school grad who’s posturing as the next Steve Jobs, when in reality he drinks bad beer at a lame, faux-hipster sports bar while watching cable televison. Or something. Dude, hop on your fixed-gear bike and standstill/peddle your sleeve tattoos out of here.

Translated: the proverbial ignite moment, that spark of innovation, is not going to come from the likes of a Cloudera or a Platfora or a (banal noun)-(o|e)ra… But it’s going to come, probably not many moons away. It will be in apps.

the sound of disruption

Thirty years ago, I went into a field called “math science”, i.e. how to build predictive analytics as software apps. Stanford — the Statistics department chairman, Bradley Efron, in particular — had put together an interdisciplinary degree which combined math, statistics, operations research, programming, engineering, etc. At the time, most of my peers in the program went on to become insurance actuaries. I went instead to do graduate work in machine learning and distributed systems.

For nearly two decades, most employers could care less about any quantitative background. They wanted C++ software engineers working all day on APIs from Sun or Microsoft or Oracle. Or they wanted managers. Then, in about 2000, came the sea change.

Right about the same time as ticker symbols for Apple and Google and Amazon were strolling up to their respective launchpads, some people began to look at my resume and ask a different line of questions.

I’ll always remember the first: a microchip vendor — one which makes electronics for several products you’ve purchased — was getting squeezed by Intel and their silicon compiler vendor. Critical features were being deprecated, specifically to put this second-tier player out of business. The company was on notice. They had to find a proverbial needle in a haystack: out of tens of thousands of circuit designs, they had to identify the 1% which would no longer be licensed — then redesign those. Quickly.

An internal team at the company had tried, but given up. Too much data for their techniques, it would’ve taken years to resolve. The company hired an electronics consulting firm in Austin, and engineers went to work, but gave up as well. Too much data, not enough signal. I got called in, as a “Whatever, just see if you can get anywhere” last-ditch effort. About 20 lines of Perl and one relatively simple equation later, I dumped my results into a scatterplot.

One of the lead circuit designers picked up my plot off the laser printer and began laughing. Loudly. The whole office heard him.

His manager grew annoyed: “What?! Why are you laughing?!”

Engineer: “He found it.”

When I turned in my invoice, the manager glared. “Look,” he said in a growl, “Just go somewhere for about three weeks. Bill us the whole time. Then come back and turn that in.”

My brows furrowed, this was a high-dollar rate for 2000.

“If you don’t pad that damn invoice…” he paused, “You’ll make both us and our customer look like complete fools. Piss a lot of people off.”

That’s the sound of Disruption.

More than a decade later, the summary graf of my resume reads like bullet points from Geoffrey Moore’s slide. Collab filters, anti-fraud classifiers, predictive analytics, etc. Even in the past few years, when HR people have read that resume, several looked up with a frown, said they thought that kind of work was better suited for business analysts — yadda, yadda, yadda, keep following the herd: you put the “botch” in “beotch”.

At a time when lots of business (start-ups as well as enterprise) are starving because they cannot hire Data Scientists, I’ve been busy building teams. Teams which delivered $MM results. I’ve hired about thirty people onto Data Science teams within the past few years — at a time when many start-ups would feel lucky to hire one. #justsayin

mal*wart

I read one of the most imbecilic essays recently from Forbes/Quora: “What Would Be The Global Impact If Wal-Mart Abruptly Shut Down?” Essentially, a hagiography stating that Wal-Mart is too big to fail, that the consequences on the US economy, the global economy, would be catastrophic. Translated: may require an enormous bailout, soon.

[In case you hadn’t guessed, I just threw up a little bit in my mouth.]

What. A. Fucking. Moron. The reality is that Wal-Mart hasn’t been doing so well over the past decade. Not if you peel back enough layers of PR. Not since they tried to bamboozle the LA city council. And failed. Moreover, folks at Amazon could really care less which Senators or SecState/former-first-lady the execs in Arkansas have in pocket. Bezos has positioned to take over 150% of Wal-Mart’s business the picosecond after Bentonville implodes. Sears and Target have reinvented themselves specifically for that very instant. So long, good riddance. Remember the point about the Global 1000 on notice? About the importance of business fundamentals?

sears, a.k.a. that web site which kinda looks like amazon

A third keynote talk that day was by the Sears CTO, Phillip Shelley. I had packed up my laptop and backpack, and was getting ready to walk out of the auditorium. After his first few sentences, I put my stuff back down and started taking notes.

Dr. Shelley mentioned how Sears started as a mail-order business a century ago, though more recently got completely kicked by another “catalog” called Amazon. Now they are leveraging Hadoop + R + Linux/Xen private cloud (srsly, this is from the Sears CTO?!?) to reinvent their business with 100x more detail on regional pricing models. Literally calculating personal pricing discounts for individuals, multiple times per day, specifically for mobile.

Sears: core algorithm moved from [6000 lines of COBOL on mainframe with 3.5 hr batch window] to [50 lines of Hadoop app on Linux with 8 min batch window], while reducing TCO for enterprise IT by two orders of magnitude. So much success, that they’ve spun it out as a new business line called MetaScale.

Brilliant strategery by Sears. Some seriously high powered Data Science talent walked out of that auditorium musing how they wished their VP Engineering was half as progressive as Sears. Srsly? Um, that’s called a PR coup. [Literally at the same moment as Wal*Mart had HR droids spamming the audience with whispers and rumors of lucrative salaries. Gak.]

emergence of the confidence economy

What’s the deal? It’s about confidence. Those giants in the Global 1000 which Geoffrey Moore says are on notice? They got that way by believing that business is largely about who barks the loudest, barks the longest, and cuts the most deals under the table. The proverbial alpha male in a wolf pack.

Wal-Mart would be a prime example, in my opinion. Their business is predicated on fundamentals which simply do not hold. Misplaced confidence. Thanks to people like Hillary Clinton, Wal-Mart has gained much influence on the House and Senate floors and the halls of the State Department and the UN assembly. In other words, so long as we manage to keep fuel costs artificially low, Wal-Mart’s market valuation will keep growing. So long as we believe that bullying vendors, conducting intelligence operations against the rest of your ecosystem, etc. — that these kinds of practices are ethical and sound in the long-run, then Wal-Mart will keep growing. Bullshit. Go look at that stock chart again. Wal-Mart is about tall white guys in dark suits, acting like complete pricks, destroying and plundering anything they can get their grimy paws on. Richard Gere in Pretty Woman, before he gets Julia Roberts. And not much more than that. On notice.

Moore is pointing out, in my opinion, that the issue at hand is about uncertainty. The point of establishing a corporate charter was always to externalize risk and perpetuate wealth for shareholders. That was true four centuries ago, when the first transnational was established, and has been true ever since. The modus of that mechanism is a process called sublation. The train wreck for sublation is uncertainty. In an environment where uncertainty holds sway, having real-time analytics from petabytes of customer data wins out over having a Senator in pocket. Any day of the week. The antidote for uncertainty is confidence. While there had been a regime of an “Attention Economy” extant for the past two decades or so, we’re now entering a new regime of the “Confidence Economy”.

Here’s the deal: people like me like those of us in Moore's lecture have been the “secret sauce” fueling the rise of Amazon, Google, Apple, etc. We use techniques which are mostly not well understood outside of Langley and the hedge funds. The tools of contemporary corporate assassins. Guys in suits who act like pricks in lieu of practicing business fundamentals — those guys are our targets. The modus is Disruption. If you have an MBA or a CxO title and not much else to back it up, I put food on my family’s table by being a sniper paid to hunt you. Lots of *great* food. And some of the best wines available. I shake the tension out of my hands, correct for wind and distance, draw a bead, take a deep breath, squeeze the trigger. Kill shot.

The challenges faced by Data Scientists are daunting. On one hand, most mathematicians lack enough solid engineering to create killer apps. Conversely, most engineers lack enough math to make any headway on the business data. Most business analysts lack enough of either the math or the engineering to be worth hiring. Data Scientists provide all three areas of expertise: the engineering and the math and the business insights to contend with mountainous torrents of data, and move the needle. On the other hand, Data Scientists must also speak truth to power. In any given business, there will be winners and losers. Executives, people accustomed to their own power, taken down. Meanwhile, we Data Scientists come prancing into a business, we do our magic, and consequently we point out which executives are bullshit and must be “executed”.

The reason why I’ve built Data teams at a time when others are starving is simple: confidence. Sure, I’ve logged three decades of machine learning, statistical modeling, data management, distributed computing, etc. When I talk with a grad student about their work, I can tell them in 25 words or less what they need to do on their first day at work to become regarded as an great asset to the team. They already know the techniques, but crave confidence. Into the trenches, fresh-out, having to speak truth to power. They’ll be placed into some faltering business unit, run some detailed analysis, and point out that the VP who’s been arguing loudly was completely wrong for the last N years and his/her ego cost the company several $MM. You can bet that those execs will return fire. However, a person like me is confident that we can get a kill shot. I show new folks how to draw a bead and squeeze the trigger. Been doing it for a long while, and will be doing for a long while more.

Snipers have an eerily pragmatic sense of confidence. And, by the way, that’s a peculiarly difficult job. Praise goes out to the men and women who serve their countries in uniform — when the cause is just. [FWIW, before tackling the challenges of data+science, I wore a military uniform and carried a rifle. Sniper training has become invaluable.]

my opinions

#1: Enterprise suffers because so many people in the corporate leadership ranks (or rather, amongst those clawing and scrambling to make their way into the corporate leadership ranks) consider themselves to be a different caste — if not a different species all together — from the rest of us who do not have a salaried position with a transnational. In a “Let them eat cake” world fraught with trillion-dollar bailouts, that’s not a particularly good way to future-proof. Moreover, this is why VCs are vital… to demolish that kind of hubris via constructive Disruption. #justsayin Word. Up. Bitches.

#2: Enterprise tooling, which is now mostly dependent on JVM-based apps, suffers because it has embraced “Convention over Configuration” … CoC has its place. I can imagine that it’s an excellent idea for heart surgeons to have a standard toolset, with scalpels in the exact same positions, etc. CoC is not a particularly good way to manage complexity and uncertainty, because it simply displaces major problems into the build system. Ultimately, it fails too much and impedes spin-up. Which, I believe, represents an enormous, ticking time bomb in Enterprise. Here’s a challenge: Show me a metric for the median period it takes in your business for a newly hired engineer to push code changes into production use which is adopted by at least 80% of your customer base. Now show me a metric for the media period it takes in your business between the point where a product manager identifies a needed feature and a newly hired engineer is ready for spin-up. From those, I’ll make a prediction based on that metric for how well your business will survive the “on notice” condition which Geoffrey Moore described. Better clues for navigating complexity and uncertainty can be found in the works of Ilya Prigogine or Stephen Wolfram. To wit, functional programming is more likely to address complexity, real complexity, and also more likely to attract top talent. CoC, not so much. Perhaps your enterprise business addresses Main Street instead of Early Adopters... Recall that Google and Amazon and Apple crossed the chasm by recruiting armies from grad students — at a time when most other people erred on the side of average joes. Remember the point about real-time analytics? It counts for training your people, then retraining, and retraining, constantly — to grapple with uncertainty. Kill shot. Global 1000.

#3: MapReduce will be unrecognizable within three years. Hadoop Summit will become something quite different after Hadoop bifurcates and gets sublated into Something Else. For example, it would be not difficult to use the Simple Workflow Service from Amazon AWS to implement MapReduce using the core part of Cascading… a different kind of MapReduce, which is not constrained by JVMs… which could scale much more gracefully and robustly… which could out-perform Google infrastructure and avoid attempting to re-create the industry in the image of Yahoo! At which point, one could deploy functional programming blocks at enterprise scale, without having to rely on the morass of enterprise build tools. Hmmm… may need to get a term sheet for that one.

Geoffrey Moore, we may have a few answers for your questions.









2006-12-20

three laws of avatarics

Susan Wu, an associate at CRV, wrote an excellent critique recently in the article Second Life: Incredible innovator, but probably not sustainable:

Their high technical barriers to participation and the fact that SL is a closed standards system ultimately deters them from reaching mass market adoption.
She followed a few days later with an example from the other end of the spectrum:
APIs, like open source, facilitate economies of scale around the development process and create network effects for the core product.
I agree with both points, but I'd like to look in closer detail about why a virtual world suffers by being "closed", and what it means to be "open". First, a little background... It's important to state: we do not live in virtual worlds, rather we project aspects of our persona into these worlds. We engage in interaction with other projected or automated personae - but we do so solely as observers of those worlds. For instance, the avatar sitting next to you at a conference in Second Life is most likely some projected persona of another observer, and the NPC you confront inside World of Warcraft is an automated persona, functioning as a relatively limited observer.

One area of systems theory called autopoiesis seeks to apply rigorous definitions from biology to explore notions of "openness" and "closure" for systems. It also examines projected aspects of observers, sustainability and regulation, and some basis for defining "cognition" in terms of interaction. I'll dive into systems theory in gory detail later, but it seems apropos for a review of Susan Wu's perspective. For now let's make a point that a virtual world can be called a "domain of discourse", and that we as users and businesses involved in a world represent projections of that world's observers.

Let's also note a special term: "personification". I use it to describe how we project our persona, our identity, crossing Edward Castronova's membrane into the many different online worlds. In the sense that "Web 2.0" successes were built on a foundation of "personalization", the successful virtual worlds will likely be those which find effective ways to manage and leverage "personification".

Some interesting conclusions follow from systems theory - issues not usually associated with "virtual reality", "social networks", or media in general, such as implications for how complex behaviors in-world may become computable and perhaps predicted based on cellular automata. More to our point, these open/closed world issues speak to what is permitted as actions by automated systems and intelligent machines. In other words, the relative "openness" of a virtual world such as SL is, according to theory, strongly coupled with what we would permit or require of robots. Not surprisingly, given how recent furor over "CopyBot" has created some of the fiercest philosophical and commercial disputes in SL's history.
***
As the Technical Director of HeadCase Humanufacturing, Inc., a start-up focused on avatars and related technologies, I recently looked online for a term to describe the study of avatars. While there may be several cheesy sites advertising "avatars" as personal icons for chat boards, there seemed to be little in the sense of science or other phenomenology. The fact that so many of us visit virtual worlds such as SL may currently be hot news in the pages of Wall Street Journal or Financial Times; however, what we do with our personifications in those worlds does not yet seem to rate its own name as a branch of academia.

Thinking that "avatars" might be considered analogous to "robots", I tried searching for the term "avatarics". Only one search hit appeared in Google, located in a German entry of Wikipedia (see English translation by its author). That page describes the concept of intellectics as suggested by Wolfgang Bibel:

Intellectics explains human intelligence by designing systems that possess it.
Bibel includes avatarics as a sub-discipline of intellectics, roughly translated as: The science that seeks to design ... anthropomorphic intelligent autonomous agents that act in virtual environments.

Now that we have a name to use, let's return to our point about virtual worlds, openness, and robots. Parallels between avatarics and robotics which Bibel mentions are not new. Many researchers - and several disciplines other than systems theory - have analyzed the overlap of robotics into virtual worlds. As virtual worlds evolve, a large existing body of engineering, sociology, economics, etc., could potentially be carried over from robotics - and vice versa. However, "what we would permit or require of robots" seems particularly interesting.

That got me thinking about Isaac Asimov, who stated his famous Three Laws of Robotics in 1941:
  1. A robot may not injure a human being or, through inaction, allow a human being to come to harm.
  2. A robot must obey orders given it by human beings except where such orders would conflict with the First Law.
  3. A robot must protect its own existence as long as such protection does not conflict with the First or Second Law.
For what it's worth, Asimov "prefixed" another law the following year: A robot may not harm humanity, or, by inaction, allow humanity to come to harm. The first three gained the most notoriety, and let's focus on those.

There have been several variations on Asimov's original theme, popping up throughout literature and science. For example, Mark Tilden took a pass at those Three Laws as:

  1. A robot must protect its existence at all costs. ("assert thyself")
  2. A robot must obtain and maintain access to its own power source. ("sustain thyself")
  3. A robot must continually search for better power sources. ("get thee thither in a whirlwind")
Tilden's view lent a more practical, "life coach" approach for earnest robots. (Elizabethan parentheticals added to track subtext).
***
I tend to agree with elements of both Bibel and Tilden. Though my university studies and 20 years since in the industry have generally focused on "AI", I rarely believe the claims attributed to Strong AI. Bibel's description of "designing systems that possess human intelligence" meshes well with my preferred approach: solutions based on autopoiesis as an area of systems that incorporates a rigorous view of cognition. In other words, I build systems as networks of human interaction, where both the humans and the computers participate together as observers.

As for Tilden's view... When we first began to write specs for HeadCase, we sought technical criteria for rating the effectiveness of the avatar API in a particular online game or virtual world. In our development work, we have found the following three criteria to be helpful in determining the "openness" of various virtual words, with respect to third-party vendors:

  1. The world must provide external API for controlling avatars.
  2. The world must allow third-party to include code plugin/module.
  3. The world must support use of external web services (HTTP, XMPP, etc.)
As a third-party vendor in virtual worlds, I enjoyed how our criteria followed Tilden's pattern of "assert", "sustain", and "thither". I offer these criteria as the Three Laws of Avatarics for evaluating the viability of a virtual world with respect to its broader, external context of social and economic systems. Consider these as preconditions for the network effects that Susan Wu mentioned.

It's clear that SL is lacking (read: broken) in terms of its avatar API and LSL scripting. Much of what could be useful has been disabled, often justified by needs for security. Competing virtual world platforms such as There.com don't even begin to meet the criteria listed above. While there is the excellent Lua scripting language and interesting API support in WoW, it appears that Blizzard has a relatively poor track record when it comes to enabling third-parties and external services (though I'd love to be proven wrong about that). Considering the wealth of "Web 2.0" resources online, both Linden and Blizzard appear to have sequestered themselves into cozy little hideaways. Is that wise?

Intuitively, the answer would seem to be a loud, resounding "No!" Much of what has succeeded on the Internet has worked because of its inherent openness and interoperability. In terms of virtual worlds, however, I wanted to find a more analytic answer.
* * *
Back to our discussion of systems theory, Maturana and Varela stated in their seminal 1980 text on cognition that, "Everything said is said by an observer." (See pp. 38-40, especially...) Considering their discourse, I have a hunch that M&V would have described a virtual world as a "domain of discourse", had such ideas been within their realm of experience; it'd be a fun question to ask Maturana personally.

Randall Whitaker has commented on their work extensively. He paraphrased about observers with:
A cognizing system engages the 'world' only in terms of the perturbations in its nervous system, which is 'operationally closed' (i.e., its transformations occur within its bounds). To the extent that the nervous system recursively interconnects its components (as in our brains), the organism is capable of generating, maintaining and re-engaging its own states as if they were literal re-presentations of external phenomena.

I wouldn't expect many SL or WoW users to endure the post-modern jargon beloved by us systems theory geeks, but the part about "generating", "maintaining", and "re-engaging" makes an interesting callback to our "assert", "sustain", and "thither" criteria for virtual worlds. Whitaker and others describe more detailed properties of complex, sustainable (read: living) systems, including "self-creation", "self-maintenance", "self-reproduction", etc. Fritjof Capra and a few other authors have even managed to craft explanations of complex systems into entertaining texts. While their descriptions may not be quite as tightly packed as the "Three Laws", it looks to me as if Tilden was headed in a good direction.

Whitaker's mention of operational closure deserves a much closer look. Consider that Maturana and Varela had been working in biology - and think about some of the properties of living cells. A cell wall prevents the all-important protoplasmic goo inside from spilling out into its environment. At the same time it allows food to be "captured" and wastes to be "expelled". The system has properties which are both "open" and "closed". In the dynamics of a living cell, this effect provides operational closure as means for using selection to determine what goes in or out... as M&V described, in a process of cognition.

Looking at computer systems, particularly looking at behaviors on large networks, we find a similar property of operational closure at the heart of enterprise network security - at least for any good security practice. This is where Linden and Blizzard miss their mark, in terms of sustainable business models. On one hand, Linden has clamped down on its API and third-party support, ostensibly for the sake of security, and yet they've left SL flying high, wide and handsome when it comes to security breaches - which have nearly become a weekly spectacle. Their operational closure has been "aimed" in the wrong directions. Blizzard, on the other hand, is simply closed - with not much hope for sustainable "protoplasmic goo" there.

In contrast, based on theory reported by Maturana, Varela, Winograd, Capra, Whitaker, et al., I have a hunch that we can determine reasonable estimators to indicate whether or not a virtual world could sustain itself in the long run. In particular, I'd point toward the (relatively obscure) works of former Stanford professor Niklaus Luhmann, who made controversial applications of M&V's autopoietic theory in sociology. Again, be forewarned about profuse jargon, dense texts, and systems theory geekness, but if you can endure it you might gain a sense for an interesting theoretical basis of where and how virtual worlds may evolve. Required reading in the study of avatarics.

***

In my survey of the emerging metaverse, I must give ample credit to Trevor Smith, who articulated a very good set of criteria for Ogoglio City - looking beyond SL toward a next-generation approach. I agree with Susan Wu that APIs, like open source, facilitate economies of scale, since the trajectory of "Web 1.0" and "Web 2.0" seem utterly contingent on that point. From that perspective I highly recommend taking a good look at Trevor's work, along with a few others who seem to get the point. Efforts which follow our criteria also include Croquet, Uni-verse, and the RESTful, Ajax-based Hive7. It looks as if Multiverse could also fit the previously mentioned criteria, though its "60% of revenue" requirement seems questionable as a sustainable practice, albeit in line with their support base in Redmond and Hollywood. It recalls the haphazard form of closure practiced by Linden and Blizzard.

Meanwhile, we're keeping a close watch on avatarics and operational closure as these virtual worlds evolve. We have good company too, considering that IBM appears to be watching for similar properties to emerge.