Bubble & Build: The 2025 MAD (Machine Learning, AI & Data) Landscape

A market map + 25 crisp ideas for an over‑heated, heads‑down year in AI & data

Anyone who has been near AI this year has felt both: the froth of record funding and the grind of real deployment. Bubbles can be silly; they also finance railroads. 2025 is the year AI shifted from chatbots to systems that actually do work: agents with tools and memory, wired into governed data, powered by reasoning models. That tension—speculation funding infrastructure—frames everything that follows.

Welcome to the 2025 MAD (Machine Learning, AI & Data) Landscape, our eleventh edition since 2012 (an almost annual effort, prior editions here). This is the biggest redraw in recent memory. We made the editorial decision to substantially cut the logo count—from a peak of 2,000+ last year to ~1,150—to make the map legible and, frankly, possible amid an explosion of new companies and products. We also gave more space to the hyperscalers and pure-play category leaders (NVIDIA, Databricks, Snowflake, OpenAI, Anthropic etc) to reflect where market share, distribution, and developer gravity sit today.

Structure-wise, we made major edits. We deleted a few sections (notably folding the former “open source” box into the broader map as open weights/OSS now permeate every layer) and added others, for example an explicit agent stack (agent platforms, agents infra/tooling) and local AI (local/on device LLM runtimes). The result is a cleaner flow from data to infra to ML/AI to agents/applications.

As in recent years, the landscape is available both as a:

With that, here’s our roundup of the 25 ideas for 2025: what’s breaking through, what’s consolidating, and where the next wave of value is likely to accrue.

(Thank you: as every year, the MAD landscape is a team effort.  Major thanks to the FirstMark crew – Aman Kabeer (co-author), Leah Levine (logos/logistics), Ben Winn (promotion) and Ryan Sullivan (blog refresh), as well as Paolo Campos (PDF design) and Jonathan Grana (interactive version, which he does as a side/promotional project at Go Fractional)

Macro & Markets

1/ Bubble dynamics without the brakes. The market is frothy again, but this isn’t 1999.   Money is plentiful and valuations often stratospheric, with a clear “AI premium”, especially around agents, frontier AI, and anything growing fast (retention concerns notwithstanding). As tends to be the case in paradigm shifts, capex/opex are front-loaded. Demand will need to materialize in a big way if the space is to stick the landing, but habits often take time to change and adoption is uneven. Add the human reality: many teams are running at 996-ish intensity—frenetic sprints that accelerate shipping velocity but raise burnout risk. On the other hand, the paradigm shift is clear, revenue is real, growth is often impressive, unit economics are legible, and the denominator is bigger.  The paradox of 2025 is that hype and fundamentals are both up; history suggests fallout can arrive before the payoffs. The year’s tone is acceleration, with nerves attached.

2/ Fragility: circularity and customer concentration. Under the big growth numbers, a lot of money is looping around a small set of players.  Some deals look circular: OpenAI inks massive GPU buys with NVIDIA while NVIDIA commits giant investment back into OpenAI; plus a multibillion chip pact with AMD that includes an option for OpenAI to buy a stake. Similar patterns extend across the stack: industry financing and supply agreements increasingly tie model labs, chipmakers, clouds and AI startups into webs of mutual dependence, prompting “roundtripping” concerns. Customer concentration makes the AI ecosystem less shock-resilient: a large share of spend runs through a short list of hyperscalers and frontier labs, while several breakout vendors lean on a handful of outsized customers – great until a growth pause or policy change ripples through. API terms, safety updates, or rate limits can impact numbers overnight. The engine is running hot, and the load is unevenly distributed.

3/ Big picture fuzzy, near-term very real: Credible voices are split on whether progress is plateauing or we’re simply missing the next exponential; paths to AGI/ASI remain undefined and the very definitions are fuzzy; even the doomer drumbeat has quieted. Meanwhile the short-to-medium term feels unmistakably concrete: an avalanche of AI slop across video, text, and code is arriving alongside ever more pressing concerns about jobs: how much changes, how fast, and for whom. The stakes are immediate even if the endgame isn’t. And, as usual, human, political, and societal responses are lagging the speed of the technology.

4/ Labs vs. incumbents: different balance sheets, same race. The fight to dominate AI is fiercer than ever, and the field is uneven. Big Tech has massive distribution, huge product suites, and budgets to bundle, wait out cycles, and grind. Google clearly regained momentum in 2025 with a string of headline AI launches, and Meta ramped ambitions with its SuperIntelligence Lab; both run on hugely profitable cores and near-infinite balance sheets. Independent frontier labs, by contrast, need step-change breakthroughs to justify valuations. New names—SSI, Thinking Machines, Reflection—joined the top tier, right as agent/reasoning heat rose (and distribution hurdles remained). OpenAI remains the clear leader and keeps raising war chests; Anthropic isn’t far behind, but how long can capital run at those levels? Users win either way: incumbents bundle, labs dazzle—while M&A stays the bridge between missing pieces and market access.

5/ IPOs & public comps: the window is open (selectively).  CoreWeave’s March debut did what the market needed: a clean AI-infra IPO that’s traded well since. Palantir is the lightning-rod comp, riding a premium EV/NTM multiple (~80–90× lately), which should embolden late-stage filers. Next: Cohere says it could IPO “soon,” Dataiku has picked banks, while Cerebras withdrew its S-1 after a fresh raise.  Meanwhile, the top 10 or so private AI players have little incentive to go public, given access to capital and strategic flexibility; but if/when Databricks (> $100B private) and the frontier labs (OpenAI, Anthropic) do eventually go out, expect bonanza, record-breaking IPOs.

6/ M&A: consolidation and talent wars. Big players tried to build full agent stacks, found it harder than it looked, then went shopping—yet even headline deals stalled (Adobe–Synthesia, SoftBank–Agility), so “buy vs. build” isn’t automatic. The wins were surgical: ServiceNow–Moveworks ($2.85B) for enterprise agents; Salesforce–Informatica (~$8B) to firm up the data control plane. Data infra is merging from within: dbt Labs and Fivetran (all-stock; ≈$600M ARR) bring ingestion and transformation under one roof. The loudest story, though, is talent and acqui-hires. Meta, in particular, went on the offense: it took ~49% of Scale AI for ~$14–15B to bring in Alexandr Wang on its superintelligence push, then set nine-figure comp markers and poached OpenAI researchers, triggering a retain-at-all-costs spiral. Bottom line: 2025 is the year of precision tuck-ins, team buys, and creative structures; true megadeals remain rare under integration risk and antitrust glare.

Research & Frontier

7/ Reasoning + RL is the frontier. The biggest leap this year wasn’t a bigger transformer; it was training models to spend compute on thinking. Reinforcement learning for reasoning—popularized by DeepSeek R1 and “o-series” style models that allocate tokens to deliberation moved the needle across math, code, and multi-step planning. Curriculum design, reward design, and tool-use feedback loops matter more than raw model size. RL is not a silver bullet —bad rewards still teach bad habits—but scaled correctly, it brings tremendous punch to pre-training. The next challenge is generalization beyond code and math to messy real-world work where “right” and “wrong” aren’t always crisp; here, richer signals matter, from business outcomes to human feedback and new benchmarks like GDPVal, which score end-to-end task chains. 

8/ Is AI slowing down? Dissent keeps us honest. Some top researchers—including guests on our MAD Podcast (Sholto Douglas, Julian Schrittwieser, Jerry Tworek)—say there’s still plenty of low-hanging fruit and years of progress ahead using the current pre-training + RL paradigm. Others urge caution: Andrej Karpathy says “agents are a decade away”; Rich Sutton’s Bitter Lesson argues that general methods plus compute beat hand-tuning; Yann LeCun pushes world models and self-supervised prediction as a different path. The debate is healthy: less leaderboard theater, more ablations, red-teaming, and real-world tasks. 

9/ Fast-moving frontiers: AI doing inventive science; robotics. We’re seeing “Move 37” ideas in the lab—models proposing non-intuitive hypotheses and paths humans wouldn’t try first. AlphaFold 3 moved into biomolecular interactions; GNoME surfaced ~2.2M plausible crystals; and Yale × Google’s Cell2Sentence-Scale 27B flagged a potential cancer-therapy pathway from single-cell data. Beyond bio, robotics is accelerating: robotics foundation models (vision-language-action policies trained on large, pooled datasets) are improving transfer across robots and tasks, while mobile manipulators log more real-world hours and autonomous lab rigs tighten the design–build–test loop. Could AI deliver Nobel-level breakthroughs, or field robots that reliably do useful work? Both feel closer each quarter.

10/ Open source (open weights) endures—through a bumpy year. DeepSeek’s R1 moment (and open-weights derivatives) set the tone, but Llama 4 underwhelmed and Meta signaled a tighter stance on permissive releases. Mistral had swings, then regained momentum; Qwen3 quietly became the “good-enough” workhorse in many stacks. On the upside, AI2 kept shipping real assets (OLMo/OLMo-2, Dolma-class data), and Reflection AI’s funding revived the “U.S. DeepSeek” narrative. Enterprises still want control and residency; startups still want margin. The future seems hybrid: route to open source when you can, spike to frontier when you must. With NPUs landing everywhere, small models will play an important role; the healthiest stacks stay plural—open and closed, cloud and device, large and small —without religious wars or vendor lock-in.

Geopolitics

11/ China assembles a parallel AI stack. China is building an end-to-end path that leans less on NVIDIA and CUDA: Huawei Ascend 910B/910C under a growing software layer (CANN, MindSpore), topped by homegrown models (DeepSeek, Qwen3, Kimi, ERNIE, etc) tuned for local data and policy. Since the DeepSeek moment, it’s been a big year for Chinese models, with Qwen and Kimi expanding in production, not just “good enough,” but competitive in several domains. Export controls slowed but didn’t stop progress; localization became a feature, signaling technical decoupling: compatible, increasingly self-reliant, and in places front-rank. 

12/ Sovereign AI goes from slogan to procurement. “Build local models on local compute” now has hardware, budgets, and real buyers behind it. The U.K. switched on Isambard-AI and finished its grid hookup; IndiaAI crossed 34,000+ GPUs and started subsidized allocations; Gulf states keep scaling national “AI factories” via G42 × Cerebras (Condor Galaxy). Europe is nurturing champions—Mistral now with ASML in its corner—while OpenAI rolled out EU/U.K. data residency to meet sovereignty asks.

13/ Energy becomes the new compute chokepoint, and nations notice. Power, not GPUs, is the new bottleneck. Datacenter location decisions now follow megawatt contracts, water rights and grid interconnects. Government court AI factories like they court fabs. Expect sovereign PPAs and nuclear/renewables co-location (Isambard-AI grid hookups in the UK; Google – TVA/Kairos SMR iplitos; Microsoft-Helion Fusion PPA). Power-first incentives will shape where models are trained and which regions win the AI buildout. Export controls still matter, but kilowatts now set the timelines.

The Business of AI

14/ Distribution beats invention (again). A whole generation of AI-native  startups is growing faster than we’ve ever seen. Products go viral on social, boards continue to fret about AI and curiosity fuels a wave of trials and tinkering. The open question is durability: true ARR or experimental revenue that churns? Incumbents often hold the distribution edge: assistants bundled with iOS/Android, Windows Copilot, Chrome, Salesforce Einstein, ServiceNow Now Assist—but not always. Partnerships and integrations can bend the curve: Cursor deepens into VS Code; Supabase rides the Lovable wave; quieter winners seep into IDEs, CRMs, and docs. Products find success by being present at creation moments (writing, coding, filing a case) where embedded beats merely “better,” and expansion follows real usage.

15/ Margins & pricing: land-grab to landing the plane. When usage climbs and customers want the newest, smartest models, costs jump fast. Hard truth: if you sit on other people’s frontier models, growth can flip you into negative gross margins—the Windsurf to Cognition story is a warning. VC money can fund the land-grab, but it won’t cover bad unit economics forever. AI startups are adapting: default to smaller, cheaper models, reserve capacity for peaks, and cache aggressively. The dominant approach is becoming price to outcomes—per case closed, per ticket resolved—with options for guaranteed throughput, so revenue tracks real results, not chatter. Winners pair cost discipline with pricing that meters actual value.

16/ Enterprise AI: deployment lags demos (but it’s landing). Enterprise deployments move slower than cool demos on social media.  Buyers want agent governance, citations, provenance, PII handling, audit trails, and tight ties to enterprise systems before switching on any level of autonomy. There is real progress on defining and implementing use cases: AI customer service, AI coding, and internal chatbots are the obvious wins; many industry- or company-specific plays will need customization, data plumbing and policy work to fully emerge. But we’re past the “Accenture phase”, and the arc is now visible—copilots to narrow agents to managed automations – and demand is firming up. Into that demand, incumbents have a distribution advantage, shipping “agent platforms” inside CRM/ERP/ITSM (Salesforce, ServiceNow, Microsoft), bundling guardrails, telemetry, and approvals in one place. But as always, never underestimate startups.  Overall, the Global 2000 enterprise market is warming up to buying and deploying AI in earnest, just not boiling yet.

AI Infrastructure 

17/ NVIDIA dominates, but diversification is real. Blackwell GB200 racks remain the reference point, yet buyers are now adding Google TPUs, AMD MI350 and, in specific footprints, Intel Gaudi 3. With rack-scale design driving TCO, many shops mix vendors for price/perf and supply, and run heterogeneous clusters under smarter schedulers—not a single-vendor monoculture. 

18/ Local AI rises: device, near-edge, and private clouds. New NPUs in laptops and phones push real work onto the device: fast, multimodal, and private by default. When tasks are too big, they spill to nearby or vendor-run “private clouds” (e.g., Apple’s Private Cloud Compute) instead of generic public endpoints. Tools like LM Studio and Ollama make local models click-to-run. On-device handles snappy UX and personal context; cloud handles heavier reasoning and shared memory. In factories, clinics, and cars, near-edge boxes protect bandwidth, privacy, and uptime. The best products hand off smoothly across device, edge, and cloud.

19/ The agentic stack becomes an infrastructure layer. Beneath apps sits a new runtime: planners and tool-calling, structured outputs and function catalogs, long-/short-term memory (vectors, graphs), sandboxed tool execution, approvals, and stateful orchestration. Around it: eval harnesses for tasks, policy/guardrails, traces and cost telemetry, dataset/version control, and rollback. What looked like “app glue” in 2024 now resembles a platform tier with its own SLAs and procurement line.

20/ Compliance, security, and red teams are foundational. Security and compliance are not a checkbox, it’s the price of running AI in production. Updated guidance (e.g., OWASP’s LLM Top 10, prompt-injection playbooks) set the bar: show where data came from, log prompts/tools/decisions, enforce policy, and prove you resist jailbreaks. Enterprises expect attestations, audit trails, and clear “break-glass” procedures, wired into the same layer as serving and storage. If it can’t be evaluated, traced, and governed, it isn’t infrastructure.

Data Infrastructure 

21/ End of an era, start of a merge. The “modern data stack” unbundling is giving way to consolidation: dbt Labs and Fivetran are combining, while platforms like Databricks keep covering the waterfront (batch and streaming, vector and graph, feature stores, governance) by equal parts build and buy. The frame shifts from “warehouse vs. lakehouse” to object storage plus open tables and a neutral catalog as the control plane. Modeling, movement, features, eval datasets, lineage, and policy are fusing with AI serving and the agentic runtime. In effect, data infrastructure and AI infrastructure are collapsing into one plane; the seams are where value leaks.

22/ Yet the data fundamentals remain more important than ever. Robust tables and catalogs, quality and lineage, and low-latency query engines have become prerequisites for agents, retrieval, and eval-first CI—not afterthoughts. Graph- and vector-augmented retrieval is moving from blog post to pattern, observability now spans prompts, tools, and cost, and compliance sits alongside performance in the same plane. The space has fresh vigor: ClickHouse’s rise in real-time analytics (now with vectors) signals demand for speed at scale, while local and edge stacks still need clean contracts back to cloud memory. Data isn’t fading; it’s been promoted to AI’s control surface.

Applications & Agents

23/ Big Labs and Platforms move up the stack. The frontier labs and incumbents aren’t content with being just model APIs: OpenAI, Anthropic, and Google/Gemini keep shipping app-layer products: voice assistants, desktop apps, team plans, and workflow builders stitched to mail, docs, and CRM. That creates platform risk and frontal competition: when the model vendor owns the surface and the bundle, it can ship into your lane tomorrow. OpenAI pushed farthest, recruiting domain experts (e.g., ex-bankers) to teach workflows, adding commerce rails inside ChatGPT, and launching a ChatGPT-first browser—while Anthropic deepened team/project flows and shipped Claude Code; Gemini tightened its consumer and Workspace surfaces. Meanwhile, models have absorbed big chunks of the “wrapper” layer: first-party structured outputs, function calling, memory, browse/code/vision/voice tools, lightweight automation, even commerce. Users get speed—capabilities landing where they already work. For startups, the wrapper cycle went thin → thick → thinner again: early UIs grew into real products (data bridges, workflow, compliance) only for the platforms to pull many features into the core. The lane that remains is narrower but real: deeply specialized workflows tied to systems of record, proprietary data/logic, and surfaces the platforms don’t, or won’t, prioritize. And for clarity: Microsoft has long lived at the app layer; the new encroachment story is led by OpenAI, with Anthropic and Gemini close behind.

24/ Vibe coding becomes the hit of 2025. Coding agents jumped from novelty to daily habit—reading repos, spinning sandboxes, planning changes, opening PRs, running tests, narrating diffs—and even “video coding” demos now show agents manipulating UIs from screencasts. Adoption has been breathtaking: Cursor and Claude Code are widely cited as among the fastest-growing dev tools ever, with reported nine-figure ARR trajectories within months. The craft shifted from autocomplete to directing and reviewing, and the stack broadened (GitHub Copilot, Sourcegraph Cody, Codeium/Windsurf, Devin, etc) toward end-to-end workflows. On the product side, Vercel v0, Lovable, and Replit turned “describe, then ship” into a production loop for tiny teams. The open question is stickiness, especially for non-professional developers, yet early cohort behavior suggests these habits may be as durable as search for coding.

25/ Modalities light up. Image, video, and voice hit a new gear: Veo3, Runway and Sora drove cinematic generation; ElevenLabs and Synthesia made high-quality voice and avatar work routine; real-time voice agents hold fluid conversations and drive tools. Vision models now parse UIs, charts, and field photos without brittle templates, and video editors jump from clips to storyboarded scenes with provenance. Meanwhile, world models—from Genie 3 to new work out of Fei-Fei Li’s group—aim to perceive and act in interactive environments, blurring creative and operational software. The bar moved from “can it caption?” to “can it perceive, plan, and act across modes, reliably?” 2026 will be a big year for modalities.

Closing thought

The 2025 MAD Landscape is a map of a market doing two things at once: bubbling and building. We redrew it to reflect reality—fewer logos, more gravity—where hyperscalers and pure-play leaders anchor the edges, agents and the data/control plane meet in the middle, and energy, not just GPUs, sets the tempo. The story lines rhyme across the map: labs climb into apps as open weights stay resilient; data and AI infra merge; enterprise deployment lags demos but is landing; coding agents become a daily habit. From here, the horizon is bigger than any single release: if we align distribution, margins, governance, and kilowatts, intelligence becomes infrastructure—and the next wave turns into compounding progress that lifts whole industries.

Quick S-1 Teardown: CoreWeave

In the tradition of previous quick S-1 teardowns (Snowflake, Palantir, Confluent, Klaviyo, Cerebras, etc), some quick notes on the CoreWeave S-1 from my colleague Aman Kabeer and I.  As in prior efforts, this is not meant to be 100% comprehensive (and it’s certainly not investment advice!).

The CoreWeave IPO is going to be fascinating to watch: partly because it is undeniably exciting, and partly because it is not going your standard tech IPO.  It  presents a profile that in some ways is typical of a hyper-growth tech unicorn (explosive growth, large losses, dual-class stock structure with Class A/B shares​, etc.), but in other ways, it is very unlike most tech IPOs of the past. Its specialized business model, heavy infrastructure focus, heavy customer concentration (Microsoft), reliance on big partners (NVIDIA), financial structure ($7.9B in debt) and unusual risk factors make it a unique case that blends the characteristics of a cloud provider, a hardware company, and a startup riding an exciting but also sudden and unproven market wave. 

SOME KEY TAKEAWAYS

Frivolity: Crypto pivot, NJ in da house: For all the jokes on social media about founders, startups and VCs pivoting from Web3 to AI, CoreWeave is an example of a business that started as a crypto mining operation, stockpiled GPUs, and pivoted to AI with spectacular success.  

In the same (frivolous) vein, for all the “all AI is in SF” mantra, CoreWeave is headquartered in… New Jersey.

The First Generative AI IPO:  Depending on pricing, there’s likely going to be tremendous interest in the IPO.  This is in part because of the well-documented dearth of tech IPOs, but most importantly  because it’s the first IPO of the Generative AI era (Cerebras, as far as we know, is still stuck in CFIUS review of their relationship with G42).  

There are very few “AI pure plays” on public markets – Palantir is one example (cue in the never ending discussion as to whether they are truly an AI company), as is (even more arguably) C3 AI.  Other than that, the way to “play” Generative AI has been to invest in Mag 7 companies.

It is no accident that the first companies to file S-1s in the Generative AI era are infrastructure companies.  The market has been forming supply-first  (chips, data centers, foundation models), with the major hope that the demand side surfaces equally meaningfully in years to come. 

Not a Real Estate Play: A negative take on CoreWeave and comparable companies one would often hear in tech circles is that the company is a “real estate play”, with limited technology and software.  The argument seemed to be supported by the fact that the co-founders of the company come from a financial, rather than technological, background.

Continue reading “Quick S-1 Teardown: CoreWeave”

AI in 2024: Never a Dull Moment

In AI, it’s been the year of big, bigger and biggest. My colleague Aman Kabeer and I just did a fun episode of The MAD Podcast where we discussed what we’re seeing in the market, our favorite trends and new stories, and where we see things going.

Here’s the video, and I summarized below a few highlights:

Continue reading “AI in 2024: Never a Dull Moment”

Is SaaS dead?

Many SaaS stocks have been getting clobbered in public markets. Some see the “end of software“.

Is SaaS dead?

What seems to be happening:

  • tough macro, cost cutting
  • AI sucking the air out of the room
  • SaaS vendors perceived as “last generation” despite best efforts to add AI quickly
  • enterprise budgets for AI are not net new, they’re taken from somewhere (SaaS budgets cut)
  • Bulk of budgets going to OpenAI/Azure etc because low hanging fruit to “do AI” (knowledge bot, coding)
  • for the more specialized enterprise apps, customers feel like they can/should “build” internally rather than “buy”

What happens next:

  • customers realize that “build” is a headache, not always the best option
  • OpenAI / Azure etc can’t / doesn’t want to build hundreds of problem specific/ vertical specific apps
  • Takes time, but legacy and new SaaS companies truly become AI-first (not just marketing), abstract away complexity of deploying LLMs
  • macro environment eventually rebounds
  • AIaaS becomes the new SaaS – what is old is new (unedited Sat morning thoughts)
  • Question is what happens to all current SaaS unicorns and public companies as this transition happens

(unedited Saturday morning thoughts)

MAD 2024: Trends in AI & Data (video)

As a companion to the 2024 MAD (ML, AI & Data) Landscape (blog post, PDF, interactive website), my colleague Aman and I had a fun chat about some key trends we see in data and AI.

Some topics we covered:

  • The impact of open source in AI
  • The future of AI agents
  • Where are we in the AI hype cycle?
  • The emerging AI stack
  • Will AI kills SaaS?
  • Is the Modern Data Stack dead?
Continue reading “MAD 2024: Trends in AI & Data (video)”

Full Steam Ahead: The 2024 MAD (Machine Learning, AI & Data) Landscape

This is our tenth annual landscape and “state of the union” of the data, analytics, machine learning and AI ecosystem.

In 10+ years covering the space, things have never been as exciting and promising as they are today.  All trends and subtrends we described over the years are coalescing: data has been digitized, in massive amounts; it can be stored, processed and analyzed fast and cheaply with modern tools; and most importantly, it can be fed to ever-more performing ML/AI models which can make sense of it, recognize patterns, make predictions based on it, and now generate text, code, images, sounds and videos.  

The MAD (ML, AI & Data) ecosystem has gone from niche and technical, to mainstream.  The paradigm shift seems to be accelerating with implications that go far beyond technical or even business matters, and impact society, geopolitics and perhaps the human condition. 

There are still many chapters to write in the multi-decade megatrend, however.  As every year, this post is an attempt at making sense of where we are currently, across products, companies and industry trends. 

Here are the prior versions: 2012, 2014, 2016, 2017, 2018, 2019 (Part I and Part II), 2020, 2021 and 2023 (Part I, Part II, Part III, Part IV).

Our team this year was Aman Kabeer and Katie Mills (FirstMark), Jonathan Grana (Go Fractional) and Paolo Campos, major thanks to all.  And a big thank you as well to CB Insights for providing the card data appearing in the interactive version. 

This annual state of the union post is organized in three parts:

  • Part: I: The landscape (PDF, Interactive version)
  • Part II: 24 themes we’re thinking about in 2024
  • Part III: Financings, M&A and IPOs 
Continue reading “Full Steam Ahead: The 2024 MAD (Machine Learning, AI & Data) Landscape”

This Week in AI: Databricks’ Acquisition of MosaicML

(This post is part of my “This Week in AI” series, which is general off-the-cuff market commentary. I’m not an investor in either MosaicML or Databricks)

$21M per employee. That’s the price Databricks is paying for MosaicML — a total of $1.3B for 62 employees (in Databricks stock, and also includes employee retention packages).

One thing is clear – if you’re going to be aggressively acquiring Generative AI startups, you’re going to have to pay up

But it may turn out to be cheap in the long term given the size of the opportunity.

That’s because, beyond any Generative AI capabilities, Databricks’ move needs to be understood in the broader context of its fierce rivalry with Snowflake.

Continue reading “This Week in AI: Databricks’ Acquisition of MosaicML”

This Week in AI: Incumbents vs Startups (Microsoft Build & Adobe Firefly)

Note: those are quick thoughts on some of last week’s most interesting news in AI. I may, or may not (!) do this on a regular basis.

This last week in AI: Adobe killed all the Generative AI design startups with Firefly, and then Microsoft killed all the other Generative AI startups with its plugins and Fabric releases.

I’m sort of kidding, but sort of not. Kidding because, again and again, founders and startups find a way. But sort of not, because the speed of deployment of AI by the Big Tech incumbents is truly something to behold. Companies like Adobe and Microsoft have, of course, massive distribution advantages. It is undeniably problematic for startups to see Adobe deploying Firefly in Photoshop and Microsoft deploying AI copilots across, well just about every single of its products for consumers, businesses and developers (see the dizzying list of announcements at Microsoft’s Build conference this week).

I don’t think, however, that the world wants a Microsoft and Google dominated AI world. The best version of the future for the Generative AI landscape is to be “polyglot” with a variety of tools and companies. Open source is going to be play a huge role and it’s comforting to see so much energy there. And I have faith startups will build the best specialized tools and vertical solutions. It’s going to be a fun ride ahead.

Here’s a good piece from Sharon Goldman at VentureBeat on some of the Microsoft Build AI announcements that triggered those thoughts above

Generative AI interview on the Infinite Machine Learning Podcast

I recently got an opportunity to chat with Prateek Joshi on Infinite Machine Learning, his excellent podcast.

It was a wide-ranging conversation about Generative AI (which I would recommend listening at 1.25x speed or more, makes me a lot more articulate). We covered a range of topics including: 

  • AI going mainstream with ChatGPT
  • The opportunity for Generative AI in the enterprise
  • Defensibility and moats of Generative AI companies
  • A mental model for thinking about what AI is best suited for, in terms of startup opportunities
  • Desirable characteristics of AI startup founding teams
  • Rapid fire: favorite books, favorite questions to ask when interviewing a candidate, why VC is a craft business

Links to the episode:

Continue reading “Generative AI interview on the Infinite Machine Learning Podcast”

MAD 2023: Top 10 Trends

Every year, as part of our MAD project, we do a presentation at Data Driven NYC about the top trends we see across data and ML/AI. (here’s the 2022 version for reference).

The presentation, done this year with my FirstMark colleague Kevin Zhang, is a whirlwind tour of top trends, as opposed to anything particularly in-depth, as we tried to keep it short. But hopefully it should provide a good overview of what’s been happening in those spaces, for anyone interested in a recap.

See below for:

  • the video (20’53”)
  • the list of top trends for easy perusal
  • the slides
Continue reading “MAD 2023: Top 10 Trends”

Interview on the Software Daily Podcast

Software Daily (aka Software Engineering Daily) has been on my podcast rotation for a while, so it was fun to get a chance to be a part of it – thanks to Jocelyn Houle who moonlights as podcast host on top of her day job at Securiti. While this was done in connection with the publication of the MAD 2023, we ended up talking a lot of about venture capital and entrepreneurship in general, including some personal stories.

The video is below, and here’s the audio-only podcast: Apple, Spotify.

MAD 2023 Interview with Ternary Data

One of the cool parts of publishing the MAD landscape every year is the conversations that come with it. Here’s a fun chat I did recently with Joe Reis and Matthew Housley, co-founders of data consulting company Ternary Data and co-authors of the O’Reilly book, Fundamentals of Data Engineering (see their recent talk at Data Driven NYC). We covered a lot of things, check it out!

The 2023 MAD (Machine Learning, Artificial Intelligence & Data) Landscape

It has been less than 18 months since we published our last MAD landscape, and it has been full of drama.

When we left, the data world was booming in the wake of the gigantic Snowflake IPO, with a whole ecosystem of startups organizing around it. 

Since then, of course, public markets crashed, a recessionary economy appeared and VC funding dried up. A whole generation of data/AI startups has had to adapt to a new reality.

Meanwhile, the last few months saw the unmistakable, exponential acceleration of Generative AI, with arguably the formation of a new mini-bubble. Beyond technological progress, it feels that AI has gone mainstream, with a broad group of non-technical people around the world now getting to experience its power firsthand.

The rise of data, ML and AI is one of the most fundamental trends in our generation. Its importance goes well beyond the purely technical, with a deep impact on society, politics, geopolitics and ethics.

Continue reading “The 2023 MAD (Machine Learning, Artificial Intelligence & Data) Landscape“

MAD 2023, PART II: FINANCINGS, M&A AND IPOs 

(note: this is part II of the 2023 MAD Landscape. The landscape PDF is here, and the interactive version is here)

“It’s been crazy out there. Venture capital has been deployed at unprecedented pace, surging 157% year-on-year globally […]. Ever higher valuations led to the creation of 136 newly-minted unicorns […] and the IPO window has been wide open, with public financings up +687%”

Well, that was…last year. Or more precisely, 15 months ago, in the MAD 2021 post, written pretty much at the top of the market, in September 2021.

Since then, of course, the long-anticipated market downturn did occur, driven by geopolitical shocks and rising inflation. Central banks started increasing interest rates, which sucked the air out of an entire world of over-inflated assets, from speculative crypto to tech stocks. Public markets tanked, the IPO window shut down, and bit by bit, the malaise trickled down to private markets – first at the growth stage, then progressively to the venture and seed markets.

We’ll talk about this new 2023 reality in the following order:

Continue reading “MAD 2023, PART II: FINANCINGS, M&A AND IPOs “

MAD 2023, PART III: TRENDS IN DATA INFRASTRUCTURE

(note: this is part III of the 2023 MAD Landscape. The landscape PDF is here, and the interactive version is here)

In the hyper-frothy environment of 2019-2021, the world of data infrastructure (nee Big Data) was one of the hottest areas for both founders and VCs.

It was dizzying and fun at the same time, and perhaps a little weird to see so much market enthusiasm for products and companies that are ultimately very technical in nature.

Regardless, as the market has cooled down, that moment is over. While good companies will continue to be created in any market cycle, and “hot” market segments will continue to pop up, the bar has certainly escalated dramatically in terms of differentiation and quality for any new data infrastructure startup to get real interest from potential customers and investors.

Here is our take on some of the key trends in the data infra market in 2023.

Continue reading “MAD 2023, PART III: TRENDS IN DATA INFRASTRUCTURE”