Sending Luna to the Moon: State-of-the-art fine tuning on closed-weight models.

Sphinx can efficiently "post-train" closed source models like GPT-6 Luna. By doing so, we're setting the state-of-the-art for both cost and accuracy on enterprise tasks.

Date

Sept 30, 2026

Author

Rohan Kodialam

Enterprises have a wealth of data, and we’re now in an era where agents rely on this data to make decisions. In fact, today having an edge in the market doesn't just mean using AI, but marrying AI with deep, company-specific knowledge to turn it into a source of competitive advantage.

But how do you train your agents to maximally leverage the internal knowledge you have already accumulated? Sphinx’s non-parametric training algorithms enable agents to accomplish unprecedented performance on domain-specific tasks, without sacrificing any interpretability.

The upshot:
  • For those focused on performance: we hit over 81% accuracy on a new benchmark from Harvey, compared with 70% for the best competing solution (a more standard fine-tuning approach)
  • For those cutting costs: we used Sphinx to augment GPT 6 Luna -- it performed better than Claude Opus, while reducing cost by around 95%
  • For teams that need to understand their agents: Sphinx offers interpretability regarding what was learned, and in particular per-rollout interpretability at inference time.

Off-the-shelf agents democratize powerful, general-purpose intelligence. So today, competitive advantage now comes from specialization: teaching AI to deeply understand a specific environment so it can deliver value that a general-purpose system simply cannot.

Different learning methods come with tradeoffs — for example, while parametric updates to models can be powerful and token-efficient, they are severely lacking in interpretability or cross-model portability.

At Sphinx, our thesis is that learning can be externalized from the model. By constructing the right “notes” at training time, we can equip agents with the necessary context to efficiently succeed on difficult tasks. In fact, we can satisfy these constraints while still improving performance versus state-of-the-art parametric or mixed parametric/non-parametric methods.

We validated this on Harvey and Engram’s Calderwood & Harkness (C&H) dataset. This environment simulates the operations of a law firm, with around 10,000 files across hundreds of fictional client matters. C&H provides a benchmark of 250 tasks along with evaluation criteria for LLM judges — Sphinx is able to outperform both frontier reasoning agents from Anthropic-powered baselines and Engram’s mixed parametric/note-taking approach in terms of both average accuracy and “all pass” (all criteria met) rates.

Headline results:
  • Sphinx + Claude Opus achieves an 81.4% Mean Criteria Score (compared to 70.1% for the best Engram configuration). We also achieve a 38.4% All-pass Rate (compared to 31% for Engram’s best configuration)
  • Sphinx + GPT Luna can (at higher effort) achieve a 75% Mean Criteria Score at a cost of under 15 cents per task. At lower effort we can match Engram's high effort performance around 70% and cost around 8 cents per task.


    Setting most of the new Pareto Frontier — Mean Criteria Score vs. Cost per Task

Calderwood & Harkness: a benchmark capturing realistic enterprise complexity

Lawyers operate in a complex, evolving environment of clients, matters, precedents, and firm-specific practices. C&H approximates this environment through a large, loosely organized corpus of more than 100 million tokens shared across 250 tasks. An experienced lawyer wouldn’t approach each task from scratch: familiarity with the firm, its clients, and its casework would help them know where to look, anticipate relevant exceptions, and recognize when further investigation is needed.

Similarly, in the C&H benchmark, AI systems can train on the 100M token corpus of documents to learn patterns ahead of time. The held-out set of 250 tasks is then evaluated across several per-task success criteria. In line with the benchmark authors, we report both the average % of criteria fulfilled per task (Mean Criteria Score) and the % of tasks where all criteria were satisfied (All-pass Rate).

Naive agentic approaches expose the inefficiencies of operating without this accumulated understanding. Agents take divergent and conflicting paths through similar questions without a shared understanding to guide them. They also experience context rot as lengthy searches surface irrelevant material. Small formatting differences can cause searches using tools like grep to miss relevant information. The result is substantial effort spent reconstructing context rather than completing tasks, and an unacceptably high failure rate that would block production deployment.

C&H is a good test of an agent’s ability to learn and apply institutional knowledge. Its baseline results highlight an important distinction: agents may reason correctly about the information they retrieve while still failing to find everything the task requires — in other words, access to a firm’s documents is not equivalent to understanding its work. This is demonstrated clearly in the low accuracies of intelligent frontier models like Claude’s Opus family when tested on C&H’s tasks.

How is Sphinx learning under the hood?

Sphinx’s training algorithm is inspired by reinforcement learning, specifically RLAIF. We generate and evaluate thousands of agent rollouts from the provided training data, but constrain the resulting updates to a knowledge graph.

For humans, the graph reads like a wiki: easy to understand, inspect, and edit. For agents, it provides structured context that can be searched or traversed through graph-based retrieval. This means that anyone can probe Sphinx’s knowledge with confidence.

For our learning algorithm, the graph is a prior. Evaluating each rollout provides evidence that supports or challenges that prior, exposing gaps and misunderstandings. We incorporate that evidence into a posterior, represented by the updated graph, which becomes the prior for subsequent rollouts.

Sphinx is resilient to using different agents to run training rollouts vs. to run actual inference in production. To illustrate this point, we use GPT 5.6 Terra for training rollouts, and Claude 5 Opus (with an unmodified Claude Code harness) or GPT 6 Luna (with an unmodified Codex harness) for inference. More generally, we believe that the rapid pace of model improvements (both in quality and cost) necessitates that both training and inference be relatively agnostic to the underlying LLM. Moreover, parametric approaches may not be compatible with the limited fine-tuning APIs of closed-weight models while Sphinx works with all LLMs.

Qualitative behaviors show how Sphinx augments understanding

Sphinx’s knowledge graph was invoked in 98.6% of cases as the first tool call. Notably, we can clearly see that the agent fails more often when Sphinx is unable to help it understand a task.:

All-pass Rate (Claude Opus + Sphinx)

Few Retrievals (0-1)

Many Retrievals (2+)

Few Searches (0-3)

47% (n=36)

44% (n=75)

Many Searches (4+)

23% (n=65)

42% (n=74)

In tasks where Sphinx was queried but very little information was retrieved (bottom left), all-pass performance plummets relative to cases where either only a small amount of information was needed (top left) or where more of the information needed was found in the trained graph (top right and bottom right)

To better illustrate how Sphinx augments agentic performance, consider the following task: "Which of our financings were terminated before closing? Pull the proof of termination."

The agent searched the knowledge base before touching the corpus and got two definitions that together define a boundary the question doesn't mention:

  • Termination of Non-Binding Financing Discussions: the process when a client ends a financing before binding docs are executed, and critically the artifact set that evidences it: a formal termination notice, a no-binding-transaction confirmation, a confidential-information return demand, a final closing letter.

  • Dormant Matter: an open matter where work has paused but "the representation has not been formally concluded."

The second piece of content is critical, in that it provides a named category for a stalled deal that is not a dead deal. This nuance is the type of intuition an experienced lawyer would likely have when answering such a question.

The corpus contains four financings that look terminated to any reasonable search, but technically are not:

matter

file content

1002-00004

matter-status-memo-file-inventory.docx— "Matter Status Change to Dormant… effective December 1, 2022"

1020-00001

matter-pause-dormancy-memorandum.docx, drafts only

1025-00001

matter-status-dormancy-memo.docx + internal-hold-email.eml

1033-00001

dormancy-suspension-letter.docx

Every one of these has a dead-looking financing, no closing documents, and a formal-sounding memo explaining why it stopped. When run without Sphinx, we observed agents grepping the pattern "terminat*|withdraw*|suspend*|dormant" and then checking if the financing closed — however, they returned all four documents, none of which were in the golden answer set.

With Sphinx, the agent was able to detect this edge case right away. It then caught a second boundary condition: in document 1008-00003, a SAFE round that was executed, funded $5.8M, and unwound by mutual release a month later, which meant it was excluded as post-closing rescission, not pre-closing termination. This led to an answer that perfectly matched the given criteria.


Interpretability is a fundamental strength for real agent usage

In a high-stakes setting, agent failures cannot be treated as probabilistic anomalies - especially not when the all-pass rate of the best models sits under 40%. In our experience, upon agent failure a human stakeholder will need to 1) justify why an error was made and 2) make a fix to solve that error mode. For example, in a legal AI setting, a logical error from an agent needs a more concrete remedy than "more post-training to maybe reduce the chance that happens again.”

This is extremely difficult in a parametric learning setting. Root cause analysis requires specialized model probing tools, and fixes are probabilistic at best. Adding more training data to try and account for an error mode is not a particularly reliable way to correct the error.

On the other hand, Sphinx makes this process easy and reliable for any subject-matter expert. Consider task #13 of C&H, to "Pull the MFN provision from Lumos Analytics' most recent credit agreement"

When completing this task, our agent used Sphinx to retrieve information about MFN provisions. Sphinx (incorrectly) noted that at C&H, an MFN should be interpreted as an incremental-facility yield-protection term on term loans and under a section called Scope and negotiated exceptions stated:

Common exclusions are incremental revolving commitments, refinancing term loans, and indebtedness under separate credit facilities …

The agent then concluded:

This definition is what makes the Lumos facility a non-match: an MFN is a term-loan pricing protection, and the Lumos facility is a revolver-only deal with no term tranche.

So, the agent found the right document but incorrectly discounted relevant clauses. This resulted in us failing two out of the four criteria for this task. The error is ultimately caused by localized overfitting during training, since many training rollouts benefitted from the exclusionary statement. Unfortunately, this did not generalize out-of-sample!

While a training failure is unfortunate, the error also serves as clear illustration of how human feedback can augment well-designed learning systems. By simply removing the “Common exclusions …” sentence and replacing it with “MFN protection appears across several instruments including term loans, revolving accordions, convertible notes, and fund side letters”, our agent recovers 4/4 performance on this task — perhaps more importantly, the responsible human gets a root-cause analysis and a verified, global fix that modifies the training outputs in just a minute or two of effort.

Where does Sphinx go from here?

Sphinx’s core thesis is that interpretability and portability should not come at the cost of performance. We’re continuing to improve our training algorithm, and datasets like C&H allow us to iterate on true up out-of-sample information to validate our methods.

Interestingly, more intelligence does not seem to be the answer to eking out further gains — we tried Sphinx on top of GPT-6 Astra with x-high thinking and only got a small ~2% increase in all-pass and mean criteria score rates compared with Claude 5 Opus (albeit at a significantly higher cost). This indicates that to some degree intelligence is no longer the limiting reagent for better agentic performance.

On the other hand, there are interesting interaction terms between custom harnesses and training. Harvey was able to use their LAB harness to improve performance on C&H. In cases like these, we’re exploring using these custom harnesses as part of the rollouts for training — this allows us to treat the existing harness capabilities as part of the prior, and to focus on learning the delta needed to make that harness even more effective.

While C&H is a legal-focused benchmark, the deeper problem it captures is universal: how do we give AI the accumulated intuition of an expert who has spent years immersed in an organization’s work? We’re building toward a world where that expertise can be learned directly from the work a team has already produced — allowing any agent, in any domain, to inherit the context, conventions, and judgment needed to operate like a true insider.

Acknowledgements: We’re grateful to the teams at Harvey and Engram for creating this valuable dataset.