Building the quality controls behind legal AI
For CTOs and product leaders running AI features, every prompt or model change raises a practical question: will it still do the job? Telos helped SimpleDocs build evaluation and version-management tools to make those decisions with recorded evidence.
The challenge
Evaluate changes across real product tasks.
SimpleDocs uses AI for different kinds of work, including review, extraction, document questions, and relationship detection. A change that helps one task may introduce problems in another. Model choice also affects latency, provider dependencies, and operating cost.
The team needed a repeatable process for testing changes against known inputs and explicit criteria, with results that could explain where a candidate performed poorly.
The work
Make prompts versioned product assets.
The application records prompt versions and supports their progression through evaluation and publication. Administrative tools give the team a place to manage versions and inspect the evidence associated with them.
Define what a useful answer must satisfy.
Evaluation criteria and test cases connect prompts with documents and expected behavior. Runs record scores and results, while thresholds identify criteria that need attention. The system distinguishes execution failures from evaluated quality failures.
Compare models against the same work.
Model-comparison capabilities let the team evaluate candidates against a shared suite. Provider configuration and adapters support changes to the underlying model infrastructure without rebuilding each product feature around a single provider.
Check readiness before publication.
Publication checks inspect test and evaluation state, including coverage of active cases. Administrative overrides exist for exceptional decisions. The system makes the evaluation state visible as part of the release workflow.
What the platform supports
- Prompt versions and administrative publishing workflows.
- Structured criteria, test documents, and evaluation cases.
- Recorded results and threshold-based quality checks.
- Comparisons across candidate models.
- Provider configuration and application-level failure handling.
The result
A repeatable way to assess changes to AI behavior.
The team can review a prompt or model candidate against explicit tasks and inspect the results before deciding how to use it. This creates a foundation for balancing quality, performance, and cost as the product and model ecosystem evolve.
Explore related projects
AI products · Microsoft Word
AI review inside Microsoft Word
Help lawyers apply review guidance, ask questions, and work through suggested edits in the document itself.
Contract operations · Document relationships
A current view of connected agreements
Link amendments, renewals, and related documents so teams can trace changes to contract terms.
How do you know an AI change is an improvement?
You need representative tasks, useful evaluation criteria, and a release process that makes the evidence visible to the people making the decision.