Resources / Research / Whitepaper

Are you training on your own work?

Firms are becoming sophisticated about preventing providers from retaining their raw information, and remain surprisingly unsophisticated about retaining the learning generated by their own employees.

Constellation · September 2026

The internet is understandably aflame after OpenAI's admission that they, “cannot rule out that de-identified data derived from [researchers'] usage of [its] products helped improve [OpenAI's] models,” in its efforts to solve the Navier–Stokes problem. The episode surfaced a question that will matter far beyond mathematics: what exactly does it mean for frontier models not to train on their customers' work?

At Constellation, we think there's an even broader question worth asking: are you training on your own work?

OpenAI subsequently said that an internal investigation determined the researchers' Codex prompts could not, in fact, have influenced the system, including through training.

Security, privacy, and appropriation

The natural response to the first question is that this is exactly what Zero Data Retention is designed to prevent. Enterprise and API customers receive guarantees that their prompts, outputs and proprietary information will not be retained or used to improve general models.

Those protections matter, but there is a subtle distinction between confidentiality of inputs and appropriability. The current enterprise debate is framed largely around a specific question: will the model provider train on our raw data?

A frontier lab does not necessarily need to retain a customer's documents verbatim to learn something economically useful from serving that customer. Raw customer data, de-identified or derived data, synthetic data, product telemetry, evaluations, and aggregate information about how customers use a system are technically and contractually distinct categories. Modern synthetic-data techniques can also recreate useful statistical properties of underlying distributions without reproducing specific customer artifacts. The relevant strategic question is therefore broader than whether a particular memo, model, or contract appears in a training set. It is: what can the supplier learn from serving us, and who captures the economic value of that learning?

This concern has become more visible as enterprises evaluate the newest frontier models. The Information recently reported that Palantir, Nvidia and Booz Allen have restricted Claude Fable for some sensitive uses because of its new retention requirements, and that some customers are seeking stronger guarantees before exposing proprietary information to the model.

At Palantir's September 2026 AIPCon, L3Harris described fine-tuning open models on its own proprietary information and operating what it termed sovereign AI. Its formulation was concise:

“We believe American defense companies should not be a vassal for frontier AI labs, handing over our data and institutional knowledge, hoping to rent back the intelligence it creates.”

Law firms are walking a similar path, and their work carries a higher standard than sensitivity: privilege. Latham & Watkins has purchased Nvidia GPU servers and is customizing open-weight models that it can operate under greater direct control, citing security, flexibility and reduced dependence on external vendors' future pricing and terms.

Knowledge that once walked out through the front door can now walk out through the API; L3Harris and others are building the framework to solve the first half of the equation. To us, the more interesting question is what happens to the knowledge that goes nowhere at all.

Watching game tape

Consider a fixed-income investor working with an agent like Claude. The agent writes an executive summary; she adds context about industry supply and demand, or comments from her sell-side analyst of choice. The agent builds a forward projection based on % y/y revenue growth; she replaces it with a nuanced revenue build-up. The agent proposes position sizing in the investment committee memo based on its training data; she updates it based on her years of accumulated judgment.

Each correction, across hundreds of investors, thousands of decisions and several years, becomes something else: a dataset describing how the institution thinks. In the loop of model output and revision, AI is making something that was historically difficult to capture increasingly legible: expert judgment itself.

A firm's historical documents are useful training material. But the more interesting asset is created after deployment: which output an employee accepts, which she rejects, which assumption she changes, where she intervenes, and most importantly, why. Those observations are labels on judgment, and at sufficient scale, they constitute a proprietary dataset describing how the institution makes decisions.

Raw traces, however, are not automatically useful training data. A user may accept an output because it is excellent, but also because the task is low stakes, because she is rushed, or because she failed to notice an error. A change may reflect institutional policy, personal preference, client instruction, or a one-off exception. The same correction can mean different things depending on who made it, under what circumstances, and with what authority.

The hard problem is therefore not simply logging interactions. It is converting those interactions into governed training examples: separating model errors from stylistic preferences, distinguishing individual judgment from firm policy, attributing corrections to the practice and seniority that produced them, preserving context and provenance, versioning rules as standards evolve, and maintaining the same privilege and information barriers that govern the underlying work.

This asset has an unusual property: it can only be accumulated prospectively. A firm can digitise its archives whenever it chooses. It cannot reconstruct why a portfolio manager rejected a projection eighteen months ago, or which three clauses a partner rewrote before a document was acceptable. Those judgments were exercised and most were destroyed within minutes. Every quarter of uninstrumented use is a quarter that cannot be recovered later, at any price.

That is the case for starting now, and it cuts against both of the usual answers. A twelve-month systems integration is twelve months of judgment thrown away. But an off-the-shelf platform, shipped to the customer to configure, captures whatever its schema was designed to capture, which is rarely the thing that distinguishes one firm from another. Generic instrumentation produces a generic corpus.

So we do both halves. Constellation's harness deploys locally and securely inside your perimeter and begins capturing from the first week, while our forward-deployed team builds the system around your workflows, your documents, and your standards. The asset accumulates while the custom build is underway, so nothing is lost waiting for it.

Return on investment

The immediate value of capturing this information does not depend on believing that every company should train or operate its own frontier model. In fact, the first economic returns may come well before any model is fine-tuned.

The first use is evaluation. A canonicalized corpus of real institutional work creates something no public benchmark can provide: an evaluation set grounded in the actual tasks, standards, errors, and judgments that matter to that organization. A firm can determine which provider performs best on its mandates, where a new model improves performance, where it regresses, and which tasks remain unsafe to automate. Instead of debating model quality based on a public benchmark designed for someone else's use case, a firm can measure performance against its own work.

That evaluation layer also creates the basis for routing. Once an institution can measure performance on its own tasks, it can determine where frontier capability is actually required. A sophisticated reasoning model may be appropriate for a complex legal interpretation or an ambiguous investment decision, while extraction, classification, document transformation, or routine analysis may be handled just as well by a much cheaper model. Without firm-specific evaluation, the safest default is often to route everything to the most capable, and most expensive, provider. With evaluation, model choice becomes an optimization problem rather than a matter of faith.

Only then comes post-training, where the evidence supports it. A firm with a sufficiently rich corpus and evaluation layer can determine whether fine-tuning an open-weight model, training firm-specific evaluators or reward models, or running controlled local inference generates a return above its cost. Some institutions may find that proprietary models materially outperform general ones on important workflows. Others may conclude that frontier models combined with strong evaluation and routing remain the superior architecture for years.

The point is not to prescribe the same technical destination for every firm. For one customer, the long-term asset may be a shared, permissioned memory layer that accumulates learned procedures, preferences, and firm rules. For another, it may be a set of post-trained models running on customer-controlled compute. A third may continue to use frontier providers almost exclusively but retain its own training corpus and evaluation layer so that providers remain interchangeable.

All three strategies depend on the same underlying asset: a structured record of how the institution's people actually perform and correct work.

Conclusion

Firms are becoming increasingly sophisticated about preventing outside providers from retaining their raw information, while remaining surprisingly unsophisticated about retaining the learning generated by their own employees.

This is the problem Constellation is beginning to work on with large law firms and financial institutions. We deploy inside existing workflows and systems, capture expert decision traces locally and securely, canonicalize them into a governed firm corpus, and build domain-specific evaluations on top. Where the evidence supports it, that corpus can then support routing, memory, firm-specific evaluators, post-training, and eventually customer-controlled models and compute.

The technical architecture will change as the underlying models improve. The economic principle is more durable.

General intelligence is increasingly something companies can rent. Institutional intelligence, the accumulated examples, decisions, corrections, and judgment that distinguish one organization from another, is something they should think carefully about owning.

For most of the twentieth century, companies worried that their most valuable knowledge would leave when their people did. AI gives them, for the first time, a plausible way to capture much more of it.

The question is whether they will.

If these problems resonate with you, security, knowledge management, frontier model cost, we would love to work with you.