Research
Harnesses vs. training - the new frontier for grounded enterprise reasoning


By studying an organization ahead of time, Sphinx achieves state-of-the-art performance across a variety of AI models and harnesses. We're able to outperform specialized enterprise harnesses like Databricks Genie by building organizational knowledge into agents directly.
Date
Oct 9, 2026
Author
Rohan Kodialam
Organizations use harnesses like Databricks’ Genie or Snowflake’s Cortex Code to connect agents to enterprise data. These harnesses provide powerful search, parsing, and automation, but expert work also depends on understanding a company’s specific context: how its data is organized, what its terminology means, and what prior work has established. That knowledge takes time to develop and is difficult to reconstruct each time a new task arrives.
Sphinx attempts to capture this knowledge ahead of time. We train on prior work to learn how to solve company-specific tasks efficiently and accurately, and that preparation gives agents a foundation they can reuse when new questions arrive, reducing the work they need to do at inference time.
On Databricks’ own OfficeQA Pro V2 benchmark, Sphinx paired with a default agent harness outperforms Genie on both accuracy and cost across all models we evaluated:
Our best configuration beats Genie’s best configuration by 5.6% at 40% lower cost.
With GPT-5.6 Luna, Sphinx reaches 48.9% accuracy at $0.17 per task. This is about one-tenth the cost to achieve the same accuracy with Genie.
These results significantly improve the tradeoff between cost and accuracy, making smaller models useful on tasks that otherwise require substantially more expensive configurations, and allowing production deployments in settings where accuracy is important.
In our comparisons with Genie, we opt to use the same models (generally from the GPT 5.6 family) that they did, to allow for a clear comparison. We find that newer models, such as GPT 6 Luna and Sol, establish an even more efficient frontier but are not directly comparable to results published by Databricks.
Note that costs in the chart above reflect inference costs only, and costs associated with document preprocessing and pre-training are excluded. Unlike Genie agents, Sphinx’s agents are not given internet access — since this benchmark has been on the open internet for several weeks, we prevent potential leakage by restricting search capabilities.
OfficeQA Pro V2 — a historically inspired enterprise scale benchmark
OfficeQA Pro V2 consists of 90 questions over roughly 120,000 pages of U.S. government financial records spanning 1793–2024. Questions require evidence from a variety of documents in disparate formats, that use varying terminology over time. The corpus contains many of the difficulties found in enterprise data: dense tables, revised figures, changing terminology, and accounting conventions that evolve over time.
Answering these questions requires understanding how the records fit together. An agent can find the right table and extract every number correctly, yet still arrive at the wrong answer because a figure was revised elsewhere or a familiar term actually refers to a different accounting measure.
Sphinx uses its study time to identify these relationships and preserve them for future work. For example, a correction may become knowledge about which figure to use next time, or an unwritten explanation scattered across annual reports becomes a reusable rule for aligning data across years.
Why does studying ahead of time allow Sphinx to perform well?
We broadly find that by capturing the subtleties of this large and complex benchmark ahead of time, we can enable agents to make the small expert-like decisions that ensure correctness. We find that specific examples clarify the mechanics of improved performance:
Panama Canal Toll Revenue
A task asked for the present value of Panama Canal non-toll receipts minus operating spending. The FY1927 report had mistakenly recorded $482,776.91 of tolls as profits. Using that year’s figures without adjustment would include revenue the question explicitly excluded.
The correction appeared in the following year’s report. It reclassified the receipts without changing their total, making it easy to miss if the agent only checked aggregate figures.
During its study, Sphinx captured the FY1928 correction and linked it to the affected FY1927 figures. With that context, the agent applied the adjustment and correctly answered −$55,649,725. Agents without Sphinx returned −$55,268,616, which is the result one would obtain by incorrectly omitting the correction.
The useful knowledge was that a later report changed how an earlier year’s receipts should be classified. Once captured, that correction could inform other analyses of the same accounts.
School funding in the 1930s
A task asked the agent to forecast FY1936 distributions to states, school funds, and forest roads using FY1914–1935 data. The reports contain several related measures: forest receipts, transfers, allocation credits, and actual spending. These differ in both amount and year. Combining similarly named rows could double-count funds or produce a forecast of the wrong series.
Across 12 fiscal-year reports, Sphinx captured the relationship between these measures. Statutory allocations use prior-year collections because Arizona and New Mexico’s school-fund entitlements must be determined before the receipts can be divided. State and road allocations describe where those collections are assigned; the clearing-account balance records what remains after transfers.
That relationship explains why annual collections, allocations, and expenditures need not match. For example, FY1914’s road allocation was $234,638.68, while reported road disbursements were $227,477.27. Both figures were valid, but only the allocation belonged in the requested series.
With Sphinx, the agent recognized the timing relationship and used credits in the year allocated, keeping them separate from collections and spending. It correctly forecast $1,186,678.
In both examples, studying the corpus ahead of time gave the agent knowledge it could apply to a new calculation. It could account for a correction or choose the right historical series without having to reconstruct the surrounding accounting context during the task.
Where do we go next?
Sphinx sits between the harness and model layers, using this preparation to guide more accurate and efficient work. The results here use Sphinx with a default harness like Codex or Claude Code. We do not have access to the exact Genie harness used in Databricks’ evaluation, so we have not yet measured the two together.
We’re continuing to explore if gains from Sphinx and from improved harnesses like Genie compound: that is, does combining Genie’s document retrieval with the knowledge Sphinx develops through prior study yield an even more efficient frontier.