Telos Labs
SimpleDocs · AI engineering

Building the quality controls behind legal AI

For CTOs and product leaders running AI features, every prompt or model change raises a practical question: will it still do the job? Telos helped SimpleDocs build evaluation and version-management tools to make those decisions with recorded evidence.

AI quality

The challenge

Evaluate changes across real product tasks.

SimpleDocs uses AI for different kinds of work, including review, extraction, document questions, and relationship detection. A change that helps one task may introduce problems in another. Model choice also affects latency, provider dependencies, and operating cost.

The team needed a repeatable process for testing changes against known inputs and explicit criteria, with results that could explain where a candidate performed poorly.

AI quality

The work

Make prompts versioned product assets.

The application records prompt versions and supports their progression through evaluation and publication. Administrative tools give the team a place to manage versions and inspect the evidence associated with them.

Define what a useful answer must satisfy.

Evaluation criteria and test cases connect prompts with documents and expected behavior. Runs record scores and results, while thresholds identify criteria that need attention. The system distinguishes execution failures from evaluated quality failures.

Compare models against the same work.

Model-comparison capabilities let the team evaluate candidates against a shared suite. Provider configuration and adapters support changes to the underlying model infrastructure without rebuilding each product feature around a single provider.

Check readiness before publication.

Publication checks inspect test and evaluation state, including coverage of active cases. Administrative overrides exist for exceptional decisions. The system makes the evaluation state visible as part of the release workflow.

AI quality

What the platform supports

  • Prompt versions and administrative publishing workflows.
  • Structured criteria, test documents, and evaluation cases.
  • Recorded results and threshold-based quality checks.
  • Comparisons across candidate models.
  • Provider configuration and application-level failure handling.
AI quality

The result

A repeatable way to assess changes to AI behavior.

The team can review a prompt or model candidate against explicit tasks and inspect the results before deciding how to use it. This creates a foundation for balancing quality, performance, and cost as the product and model ecosystem evolve.

How do you know an AI change is an improvement?

You need representative tasks, useful evaluation criteria, and a release process that makes the evidence visible to the people making the decision.