Client Story

Finding Price Increases Buried in 2,000 Contracts with AI

A $20 million service book with no index

Our client services medical imaging equipment for hospitals, surgical centers, specialty clinics, and government facilities, with service contracts worth about $20 million a year.

These contracts were stored on a shared drive, each with no naming convention or index. Some were searchable PDFs, but others were unsearchable scanned images or handwritten pages. Many included clauses allowing for renewal price raises of 3% or 5%, some as high as 9% per cycle. Unfortunately, the staff couldn’t easily find those clauses, so reps renewed from memory. Contracts rolled over at the old rate year after year. Even a simple question, such as which agreements expire in the next 90 days, required a contracts expert to open files one by one, which took anywhere from half a day to several days. New hires had almost no way forward.

An AI prototype built on real contracts

Our client’s leadership team wanted to understand how AI could improve contract management, and our AWS Rapid GenAI assessment was a perfect fit. After extensively discussing the company’s specific needs, we built an AI prototype that leveraged a knowledge base in Amazon Quick Suite against their full contract library.

Our executive read-out and demo to leadership quickly proved its worth. Every question they asked delivered an answer from their own contracts and linked it back to the source PDF, down to specific clients and escalation terms. We discussed their security concerns immediately, explaining how our AWS-based setup keeps contract data inside a closed system.

AI prototypes like these make AI approachable for our clients. We deliver a working implementation that has undergone model testing and shows monthly costs before asking clients to commit to a development budget.

Built to stay accurate at scale

In the production phase, we will add an automated pipeline that pulls contracts from the shared drive, a database of contract terms such as renewal dates and escalation clauses, a secured production agent, and an accuracy dashboard.

Working with 2,000 contracts makes human spot-checks impractical, arguably impossible, so we will build a set of contracts with verified ground-truth answers and test every change against it. Staff can then review flagged or low-confidence results.

A partnership that works

When it comes to AI, we’d rather prove it than pitch it. Our team can deliver a tested, working prototype and its costs before we start a full build. Clients can watch an agent answer questions about their own data to see how the systems work before committing to a budget.

Clients benefit from our pre-build model testing, which lets us evaluate different model types and costs. In this case, we determined that the lightest one, Amazon Nova 2 Lite, gave the best answers at roughly a tenth of the cost per question. We also walk clients through how their data stays protected, how it remains inside a closed AWS environment, and none of it is used to train AI models.

Under the hood

In Production, the AI agent runs on Amazon Bedrock AgentCore and selects one of four tools for each question:

Semantic search across all contracts
A deep dive into one named contract
A structured metadata query, or
A contract summary

Semantic search matches meaning instead of exact words, using vector embeddings (numeric representations of what a passage says) stored in an Amazon Bedrock Knowledge Base.

Semantic search returns the closest matches but can’t guarantee it found them all, so questions that need every result, such as “every contract with a minimum 5% escalation clause,” go to the metadata database instead. The system can read scanned and handwritten contracts and turns the words on a scanned page into searchable text. When a page is too messy for that, it uses image recognition to understand what’s there.

We scored Claude Sonnet 4.5, Claude Opus 4.5, Nova 2 Lite, and Nova Pro on faithfulness, helpfulness, output quality, and ROUGE, a recall-based text-scoring metric using 13 realistic contract questions. Nova 2 Lite led on every measure; the heavier models tended to overelaborate on focused extraction tasks.

The AI models run on private copies inside AWS, so contract data goes in but never flows back out to the companies that make the models. None of it is used to train models, either. Amazon Bedrock adds built-in guardrails that block hidden malicious instructions (known as prompt injection) and attempts to pull data out of the system. And because the agent only runs when someone asks a question, computing costs stay under $5 a month.

Your AI Prototype

Our team can take the same prototype approach with your data so you can see the impact AI can make before committing to a build. Contact our AI team to discuss what an AI prototype looks like for your business.

Scroll to Top