Why Agentic AI Is Breaking Lakehouse-Only Data Platforms

The News

SAP has acquired Dremio, the lakehouse query engine vendor that built its product around a centralized data lake architecture. Writing in a public analysis, Starburst Head of Product Emma Tippet frames the deal not as an isolated M&A event but as a symptom of a structural mismatch between lakehouse-only query engines and the demands of agentic AI workloads. Her argument is that AI agents, which reason across data estates dynamically and in parallel, expose the core limitations of engines designed for predictable, precomputed queries against a single consolidated store. The piece is authored by a Starburst executive and should be read with that competitive context in mind.

Analyst Take

The Dremio-SAP deal is getting covered as a consolidation story. That framing misses the point. The real story is architectural: a class of data platform built on the assumption that workloads are predictable and data can be centralized is now being stress-tested by workloads that are neither. Dremio’s sale is the most visible evidence so far, but the pressure is industry-wide.

Why agentic AI breaks the precompute model

Lakehouse engines historically earned their speed through materialized views and scheduled refreshes. The assumption baked into that model is that someone, somewhere, knows what questions will be asked. Dashboards are a perfect fit. Agentic AI is the opposite of a fit. An agent investigating a revenue anomaly doesn’t run a predetermined query. It runs a query, reads the result, generates a follow-up based on what it just found, and repeats. The query shape at step three was unknowable at step one. An architecture that can only go fast on questions it anticipated is structurally mismatched to a workflow that exists precisely because no one anticipated the questions.

This is not a performance gap that faster hardware closes. It’s a design gap. And it shows up in a place enterprises care about deeply: the reliability of the agent’s output. A human analyst working from incomplete data instinctively hedges. An agent doesn’t. It reasons over what it can reach and reports with equivalent confidence regardless of whether that picture is complete or partial. A gap in data access doesn’t produce a cautious, partial answer. It produces a confident, wrong one. That failure mode is harder to detect than latency and significantly more damaging to enterprise trust in agentic systems.

The five requirements AI is actually imposing on data platforms

Tippet’s piece outlines what agentic workloads actually demand, and it’s worth treating this as a vendor evaluation checklist rather than a marketing argument. The requirements are: near-total data reach across databases, warehouses, object stores, and SaaS systems; concurrency headroom that goes well beyond what dashboard traffic ever generated; context scoped to the question rather than the full catalog; freshness measured in seconds, not overnight batch cycles; and governance enforced at query time so two agents hitting two systems don’t derive contradictory answers from the same underlying facts.

That last point deserves particular attention from ITDMs. Semantic inconsistency across agents is not a data quality problem in the traditional sense. It’s a trust problem. If a procurement agent and a finance agent reach different conclusions because they hit different systems with different business logic applied, the organization doesn’t get two data points to triangulate. It gets two confident answers and no way to know which one to act on. That’s the kind of outcome that causes AI programs to stall at the governance review stage rather than reaching production scale.

What this means for the vendor landscape

Starburst’s competitive position is obvious here: Tippet is arguing, with transparency, that federation beats centralization as an architectural posture for agentic AI. The argument is coherent. But buyers should evaluate it against their own estate, not against Dremio specifically. The relevant question is whether your current data platform can reach the systems your agents will actually query mid-reasoning, not whether it scores well on a benchmark designed for analytical queries.

ECI Research’s 2026 Application Development survey found that 65.2% of respondents said 0–20% of engineering time is spent on net-new innovation. That figure reflects a real constraint: most engineering capacity is absorbed by maintenance, integration work, and keeping existing systems running. A data platform architecture that requires significant up-front ETL, schema harmonization, or data movement before it can serve an agentic workload adds to that burden. Federation, in principle, reduces it by querying systems where they already live. Whether any specific vendor delivers on that principle at production scale is a separate evaluation. Separately, ECI Research’s 2026 DevSecOps & AppSec survey found that 29.1% of respondents selected “AI-generated package risk” as their biggest open-source security concern in 2026, a signal that AI-adjacent infrastructure is under active security scrutiny. Data platforms that can enforce policy at query time, as Tippet describes, are better positioned to satisfy that scrutiny than those relying on perimeter controls applied after the fact.

Looking Ahead

The Dremio acquisition will not be the last consolidation in the lakehouse engine category. Vendors whose core differentiation was query performance within a single lake are now competing in a market that is actively redefining what “performance” means. The shift from optimizing known queries to handling unknown ones at scale is not a roadmap item. It requires architectural choices made years ago, and those choices are now either assets or liabilities. Expect more M&A, more pivots, and more vendors rebranding their existing capabilities as “agentic-ready” without the underlying architecture to back the claim.

For ITDMs, the immediate action is a capabilities audit against the five requirements Tippet names, applied to your current platform without vendor input. For engineering teams, the architectural question is whether your data access layer is designed for the query shapes your agents will generate versus the query shapes your analysts historically ran. Those are different problems, and treating them the same way is how AI programs accumulate technical debt before the first production deployment ships.

Authors

  • Paul Nashawaty

    Paul Nashawaty, Practice Leader and Lead Principal Analyst, specializes in application modernization across build, release and operations. With a wealth of expertise in digital transformation initiatives spanning front-end and back-end systems, he also possesses comprehensive knowledge of the underlying infrastructure ecosystem crucial for supporting modernization endeavors. With over 25 years of experience, Paul has a proven track record in implementing effective go-to-market strategies, including the identification of new market channels, the growth and cultivation of partner ecosystems, and the successful execution of strategic plans resulting in positive business outcomes for his clients.

    View all posts
  • With over 15 years of hands-on experience in operations roles across legal, financial, and technology sectors, Sam Weston brings deep expertise in the systems that power modern enterprises such as ERP, CRM, HCM, CX, and beyond. Her career has spanned the full spectrum of enterprise applications, from optimizing business processes and managing platforms to leading digital transformation initiatives.

    Sam has transitioned her expertise into the analyst arena, focusing on enterprise applications and the evolving role they play in business productivity and transformation. She provides independent insights that bridge technology capabilities with business outcomes, helping organizations and vendors alike navigate a changing enterprise software landscape.

    View all posts