Databricks Launches Faster Adaptive Search Model for Multi-Step AI Agent Queries
Databricks has unveiled an upgraded version of its Adaptive Instructed-Retriever that accelerates response times for AI agents needing to pull information from multiple sources, delivering twice the speed of competing models from Anthropic, OpenAI and DeepSeek.

Databricks Inc. has rolled out an enhanced iteration of its Adaptive Instructed-Retriever search model designed to reduce latency when artificial intelligence agents must perform sequential retrieval operations across multiple data sources.
According to the company, this retrieval component powers Genie Code, Genie One and Genie Agents while delivering comparable retrieval performance to major competing models—yet operates at double the speed of Anthropic PBC's Claude Sonnet 5, OpenAI Group PBC's GPT-5.6 Luna and Hangzhou DeepSeek Artificial Intelligence Co. Ltd.'s V4-Flash. Testing relied on a combination of seven internal and external benchmarks covering diverse domains and retrieval complexity levels.
The new model extends Instructed-Retriever-1, which debuted earlier this year with a focus on enhancing retrieval-augmented generation by preserving user instructions, examples and data-source schemas throughout both retrieval and response-generation phases. Databricks had previously demonstrated that this architecture achieved over 70% performance gains relative to conventional RAG on enterprise question-answering evaluations.
While its predecessor executed parallel single-step searches suitable for straightforward requests, Adaptive Instructed-Retriever targets complex queries demanding evidence synthesis from multiple locations or dependent on intermediate findings. Such multi-hop scenarios ordinarily require agents to reformulate queries iteratively, with each additional retrieval cycle raising both response time and computational overhead.
The updated model tackles this tension through dynamic optimization. Developers specify an upper limit on sequential search iterations, and the model independently calculates the necessary count per query. It halts when sufficient evidence exists and continues when additional retrieval cycles promise meaningful improvements to the final answer.
This capability serves data agents operating across expansive, frequently-updated repositories. Simple queries should complete rapidly, whereas complex discovery operations may justify extended processing and resource consumption. Databricks characterized the goal as enabling search intelligence that locates information "without wasting turns on brute-force exploration."
Training and Optimization
Databricks constructed the model using synthetic enterprise retrieval scenarios and multi-hop question sets, leveraging existing Instructed-Retriever-1 data to preserve single-step performance. The team then employed online reinforcement learning via Clipped Importance Sampling Policy Optimization, with a reward mechanism that incentivizes accurate search paths while penalizing superfluous steps lacking corresponding quality improvements.
By adjusting penalty magnitude, Databricks generated multiple model versions occupying different points along the quality-latency spectrum. Steeper penalties encourage fewer iterations and quicker responses; gentler penalties permit extended searching for enhanced retrieval accuracy. Organizations can select a checkpoint aligned with their use case, whether interactive applications or asynchronous batch processing.
Performance Example
In a demonstration case, Adaptive Instructed-Retriever confirmed that a corporation had not explicitly disclosed restructuring expenses in its 2022 fiscal income statement using two search iterations. Claude Sonnet 5 achieved identical recall in three steps, while GPT-5.6 Luna required four.
Databricks conducted these comparisons internally using proprietary benchmarks, meaning the results remain unvalidated by independent parties. The announcement omitted specifics regarding model parameter count, cost structure or availability timeline.
The system was developed with Databricks AI Runtime, a platform customers can leverage to customize models for their proprietary data and performance targets. Databricks emphasized that compact, domain-focused models can rival state-of-the-art systems by mastering not merely how to retrieve information, but when the investment in additional retrieval cycles justifies the resulting delay.

