Databricks Lakehouse Platform for Real-Time, Agentic Data and AI Workloads
Databricks is pushing the lakehouse beyond analytics: federated and streaming data, transactional serving, AI Search and agents can share one governed data foundation. What that changes for data and AI engineers, and where the limits are.

For years, data engineering has often meant connecting systems that were never designed to work together.
ERP data lives in one place. Manufacturing data lives somewhere else. Application events arrive through streaming systems. Operational databases serve applications. Vector databases power AI search. Machine-learning systems maintain their own features and serving layers.
Then someone asks a deceptively simple question:
“Can we connect all of this and let an AI agent actually use it?”
That question is where the modern lakehouse architecture gets interesting.
Databricks is increasingly positioning its Data and AI Platform as a common foundation for analytical, transactional, real-time, and agentic workloads.
The idea isn’t that every byte of data has to be physically moved into one database.
Instead, the platform combines data that is copied into the lakehouse with data that can remain in its source system and still be queried through capabilities such as Lakehouse Federation and Open Sharing, while Unity Catalog provides a common governance layer.
On top of that foundation sit Lakeflow for data engineering, low-latency ingestion through technologies such as Zerobus, Lakebase for operational serving, AI Search for retrieval, and agentic capabilities such as Genie and Agent Bricks.
For data and AI engineers, the interesting question isn’t whether Databricks has all of these products.
It’s whether bringing them together actually changes how we build production AI systems.
I think it does.

The whole platform on one page: ingest and connect, govern and prepare, serve and search, then build agentic applications on the same governed foundation.
Why the lakehouse matters to data engineers
Consider a manufacturing quality investigation.
A quality engineer might need to correlate:
- MES records
- machine and process-historian data
- supplier information
- QMS history
- 8D corrective-action records
- logistics data
- plant-level operational information
The question might be:
Which supplier lot reached the products affected by this defect, and has the same problem appeared before?
Databricks points out that questions like this cross organizational and system boundaries. Historically, answering them can mean tickets, exports, specialist knowledge, and manual reconciliation.
The architectural alternative is to make the relevant data accessible through a governed platform.
With Lakehouse Federation and zero-copy Open Sharing, data doesn’t necessarily need to be copied into another ETL pipeline simply because a new analytical question appeared. Data can remain in its source system while being queried through Databricks, while other data can be mirrored into the lakehouse when that makes more sense.
That changes the role of the data engineer.
Instead of continuously building point-to-point pipelines for every new question, more of the engineering effort can move toward:
Source systems
│
▼
Governed data
│
├── Analytics
├── BI
├── ML
├── AI Search
└── Agents
The goal isn’t “no ETL.”
The goal is less unnecessary data movement and fewer disconnected copies created solely to answer another question.
The same fragmentation problem exists in AI systems
AI systems make the problem even harder.
An agent may need access to:
Customer data
+
Operational data
+
Product catalog
+
Documents
+
Real-time events
+
Vector embeddings
+
Business rules
+
Model predictions
If every AI application builds its own copy of these assets, organizations quickly end up with another version of the data-platform problem.
One team maintains a warehouse.
Another maintains a feature store.
Another maintains a vector database.
Another maintains an operational database.
Another builds an agent-specific retrieval pipeline.
The result can look like:
┌── Data warehouse
│
├── Data lake
│
Applications ───────┼── Feature store
│
├── Vector DB
│
├── Operational DB
│
└── Agent / RAG stack
Every arrow between those systems is another integration to maintain.
Databricks’ lakehouse approach is essentially an attempt to reduce that fragmentation.
What the Databricks platform actually brings together
The platform is broader than the traditional “Spark + Delta Lake” mental model.
The current architecture combines several layers.
Data ingestion
Databricks provides multiple ingestion paths, including:
- Zerobus Ingest for low-latency streaming
- Lakeflow Connect for managed connectors
- Structured Streaming
- Auto Loader
- change-data-capture workflows
The manufacturing architecture described by Databricks combines these capabilities with Lakeflow and governed Delta tables.
Data refinement
The familiar Bronze → Silver → Gold pattern still has a role.
Raw events
│
▼
Bronze
│
▼
Silver
│
▼
Gold
Lakeflow can build, schedule, and monitor the pipelines that transform raw inputs into trusted, analysis-ready datasets.
Governance
Unity Catalog acts as the common governance layer across mirrored and federated data.
Databricks describes it as providing a common permission model, lineage, and discovery across data, models, and AI assets.
AI and agents
On top of that governed data foundation, Databricks provides:
- Genie One
- Agent Bricks
- Genie App Builder
- AI Search
- model serving
- Lakebase
The important architectural idea is that agents aren’t necessarily operating against a completely separate copy of enterprise data.
They can operate against the same governed data foundation used by analytics and other applications.
Real-time data changes the architecture
Batch processing is relatively straightforward.
You can ingest data overnight, transform it, build features, train a model, and refresh recommendations.
Real-time applications are different.
Imagine an e-commerce application where a customer is browsing products.
The system needs to combine:
Long-term preferences
+
Historical behavior
+
Current session
+
Product catalog
+
Inventory
+
Business rules
The recommendation has to be generated while the user is waiting.
Databricks published a reference architecture for a large fashion e-commerce platform in Asia serving more than 1 million monthly active users and a catalog exceeding 100,000 SKUs. The architecture combines Zerobus Ingest, Lakebase, AI Search, Model Serving, and MLflow.
The system processes approximately 1,000 events per second.
But the most interesting part isn’t the number.
It’s the split between batch and real-time paths.
Two serving paths beat pretending everything is real-time
The Databricks reference architecture separates predictable recommendation workloads from session-aware workloads.
Path A: pre-computed recommendations
For predictable surfaces such as:
- home-page recommendations
- category pages
- email campaigns
- push notifications
the system can calculate recommendations ahead of time.
The resulting ranked lists are stored in Lakebase.
At request time, the application can perform a simple lookup:
user_id + surface
│
▼
Lakebase
│
▼
Pre-computed recommendations
No vector search or model inference is required for that request.
This is an important design principle:
Don’t use real-time inference when pre-computation is good enough.
Path B: real-time, session-aware inference
Other experiences are inherently dynamic.
For example:
“Show me products similar to what I’m looking at right now.”
The user’s recent actions matter.
The architecture therefore sends current session signals directly to the model-serving endpoint as part of the request. Those signals don’t have to pass through the lakehouse first.
The real-time path can then perform:
Current session
│
▼
Query embedding
│
▼
AI Search
│
▼
Candidate products
│
▼
Lakebase feature lookup
│
▼
LightGBM scoring
│
▼
Business rules
│
▼
Final recommendations
Databricks describes implementing this multi-stage pipeline inside a single MLflow PyFunc predict() call.
That’s a much cleaner serving contract:
Application
│
▼
One endpoint
│
▼
Complete recommendation pipeline
rather than forcing the application to orchestrate every internal step itself.
Graceful degradation matters just as much
Real-time systems eventually encounter latency spikes.
If the real-time path exceeds its latency budget, the Databricks reference architecture falls back to cached recommendations or popular items.
That’s a pattern I would recommend well beyond recommendation systems.
Request
│
▼
Real-time path
│
┌───────┴───────┐
│ │
Fast enough Too slow
│ │
▼ ▼
Fresh result Cached result
For production AI, degraded-but-useful is usually better than unavailable.
Where Databricks AI Search fits
AI Search is another important piece of this architecture.
At a high level, it provides approximate nearest-neighbor vector search over data managed through Databricks.
It can also support hybrid retrieval by combining similarity search with keyword search. Databricks documents BM25 for keyword relevance and Reciprocal Rank Fusion for combining keyword and similarity results.
That makes it useful for workloads such as:
User query
│
▼
Embedding
│
▼
Vector similarity
+
Keyword relevance
+
Metadata filters
│
▼
Relevant candidates
│
▼
Agent / model
This is especially relevant for agentic applications.
An agent doesn’t just need a language model.
It needs retrieval.
And retrieval needs to be connected to the same governed data environment where the underlying information lives.
But AI Search isn’t magic
This is where the architecture needs a little skepticism.
AI Search uses approximate nearest-neighbor retrieval rather than exact nearest-neighbor search.
That means retrieval quality depends on the index and search configuration.
Databricks currently documents:
- standard and storage-optimized endpoints
- vector and hybrid search
- managed or self-managed embeddings
- full-text search as a beta capability
- endpoint and index capacity limits
- query-size limits
- application-level ACL workarounds where row/column permissions aren’t supported directly
The current hybrid search implementation uses Reciprocal Rank Fusion with an rrf_param of 60, based on the literature.
That doesn’t mean the search system is inadequate.
It means that retrieval quality remains an engineering problem.
You still need to evaluate:
Recall
Precision
Latency
Freshness
Filtering
Embedding quality
Ranking quality
Cost
An agent is only as good as the information it retrieves.
The embedding lifecycle matters too
Another detail that can easily get overlooked is how embeddings are managed.
Databricks AI Search supports multiple approaches, including:
- Databricks-computed embeddings
- Self-managed embeddings
- Direct Vector Access
- Dedicated full-text search on storage-optimized endpoints
The choice affects how the index is updated and maintained. For example, Databricks notes that a self-managed embedding index cannot simply be converted into a Databricks-managed embedding index later; changing approaches requires creating a new index and recomputing embeddings.
That’s exactly the sort of decision that should be made before an AI Search deployment becomes a production dependency.
The reference architecture I would take away
Putting the pieces together, the broader Databricks architecture looks something like this:
SOURCE SYSTEMS
│
┌─────────────────┼─────────────────┐
│ │ │
ERP / MES Apps / Events Documents
│ │ │
└─────────────────┼─────────────────┘
│
┌────────────▼────────────┐
│ Ingestion / Federation │
│ Lakeflow · Zerobus │
│ Open Sharing │
└────────────┬────────────┘
│
┌────────────▼────────────┐
│ Governed data │
│ Unity Catalog │
└────────────┬────────────┘
│
┌────────────────┼────────────────┐
│ │ │
Analytics ML AI Search
│ │ │
│ Lakebase │
│ │ │
└────────────────┼────────────────┘
│
Model Serving
│
┌──────▼──────┐
│ Agents │
│ Genie · │
│ Agent Bricks│
└─────────────┘
The important part isn’t the individual boxes.
It’s the shared governed foundation underneath them.
What data engineers should actually do
If you’re evaluating this architecture, I wouldn’t start by asking:
“Should we move everything to Databricks?”
I’d start with four smaller questions.
1. Which data actually needs to move?
Use federation when data should remain in its source.
Mirror data when the lakehouse provides a meaningful advantage for:
- analytics
- transformation
- ML
- historical analysis
- downstream AI workloads
Databricks itself describes this as a pool-or-federate approach rather than a mandatory migration.
2. Which workloads actually need real-time inference?
Don’t make everything real-time.
Use:
Batch
→ predictable workloads
Real-time
→ session-sensitive workloads
The retail reference architecture is a good example of this split.
3. Where should the semantic layer live?
This may be the most important question for agentic AI.
If humans and agents use different definitions of:
revenue
customer
active user
inventory
conversion
supplier risk
then the platform isn’t really unified.
The semantic layer needs to become a shared contract between:
BI
Analytics
ML
Applications
Agents
That’s one reason Databricks emphasizes governed semantics and natural-language access in its manufacturing architecture.
4. How will you measure the AI system?
Don’t stop at:
“The model works.”
Measure:
Retrieval recall
Answer quality
Agent success rate
Latency
Freshness
Cost/request
Cost/completed task
Fallback rate
Data-quality failures
Cost per completed task deserves special attention for agents: as we saw with GPT-6.1 Sol on Amazon Bedrock, the cheapest model call isn’t always the cheapest finished task.
The platform makes the architecture easier to assemble.
It doesn’t eliminate the need to measure whether the architecture is actually working.
What this means for AI agents
This is where I think the lakehouse becomes especially interesting.
A traditional RAG architecture often looks like:
Documents
│
▼
Chunking
│
▼
Embeddings
│
▼
Vector database
│
▼
LLM
A production enterprise agent needs much more:
Enterprise agent
│
┌────────────────┼────────────────┐
│ │ │
Search Structured Actions
│ │ │
AI Search Lakehouse APIs / Tools
│ │ │
└────────────────┼────────────────┘
│
Governed context
│
▼
Agent
This is why the lakehouse story is becoming relevant to agentic AI.
The hard problem isn’t merely calling an LLM.
The hard problem is giving the LLM trustworthy access to the right enterprise information and actions.
The limitations are worth taking seriously
The unified-platform story is compelling, but it doesn’t make the engineering problems disappear.
AI Search still has documented limits around index capacity, query size, result size, permissions, and endpoint configuration. For example, standard endpoints support high-QPS configurations, while storage-optimized endpoints trade somewhat higher query latency for much larger index capacity and faster indexing.
There are also architectural questions that organizations should benchmark themselves rather than assume away:
- How does search latency behave under your actual concurrency?
- How quickly must indexes reflect new data?
- What happens when an upstream streaming source falls behind?
- How much does the complete ingestion + storage + search + serving stack cost?
- What is the failure strategy when one component becomes unavailable?
- How will you handle embedding-model changes?
- How will access-control changes propagate through agent-facing applications?
- Where should data remain federated versus physically copied?
Those questions are more important than whether a platform diagram contains fewer boxes.
A unified architecture is valuable only when it produces simpler operations without compromising reliability, governance, or performance.
The engineering takeaway
The biggest shift here isn’t that Databricks has added more AI features to the lakehouse.
It is that the boundary between data engineering and AI engineering is becoming much thinner.
A modern AI application increasingly needs:
Data ingestion
+
Data quality
+
Governance
+
Semantic modeling
+
Real-time processing
+
Feature serving
+
Retrieval
+
Model serving
+
Agent orchestration
Historically, those capabilities often lived in separate systems.
The Databricks approach is to make them components of one governed platform.
That doesn’t mean every company should immediately replace every existing system with Databricks.
It means there is now a credible architecture where the same governed data foundation can support:
analytics → ML → real-time applications → AI Search → agents.
For data engineers, that means fewer arbitrary boundaries between the systems we build.
For AI engineers, it means agents can move closer to the actual operational data they need to reason about.
And for both, it raises a more interesting architectural question:
What if the data platform itself becomes part of the AI system?
That’s the direction worth watching.
Key takeaways
- Lakehouse architecture is expanding beyond analytics into transactional, real-time, search, ML, and agentic workloads.
- Lakehouse Federation and Open Sharing can reduce unnecessary data copies, while allowing data to remain in source systems where appropriate.
- Unity Catalog provides the common governance layer across mirrored and federated data.
- Lakeflow and Zerobus support the ingestion and refinement layer, including low-latency streaming workloads.
- Lakebase, AI Search, and Model Serving can be combined for real-time applications without forcing every request through a batch pipeline.
- AI Search supports vector and hybrid retrieval, but retrieval quality, latency, permissions, and capacity remain engineering concerns.
- The strongest architecture is not necessarily the one with the fewest systems. It is the one that minimizes unnecessary movement and duplication while preserving governance, reliability, and performance.
- The real opportunity is convergence: one governed data foundation serving analytics, ML, applications, and agents.
Sources
- Databricks Blog: Manufacturing data and AI: Connecting the product value chain (September 28, 2026)
- Databricks Blog: Real-Time Retail Intelligence: Building E-Commerce Recommendations with Lakebase and AI Search on Databricks (October 2, 2026)
- Databricks Documentation: Databricks AI Search (AWS)
Discussion
Join the discussion: create a free account or sign in to comment.
No comments yet.