Client Story

On-premises AI Answers a College District’s Cost and Policy Questions in Seconds

Our AI team worked with a large California community college district’s staff to build an AI solution that delivers faster answers about technology costs and union rules without data leaving its hardware. We worked directly with Apple and Claris to build a private AI assistant inside the district’s FileMaker system, and most questions now get an answer in two to four seconds.

Questions the district can’t afford to get wrong

The district buys and replaces technology at scale and must decide how much of its device fleet to replace, and when. And with a union-heavy workforce, a wrong answer to an everyday policy question can become a contractual problem.

Weeks of manual math

The full cost of owning and running devices and labs for the district is called the Total cost of ownership (TCO). Calculating that number takes weeks of manual work, so it rarely got refreshed, forcing replacement decisions to rely on outdated information.

Getting answers to policy questions was slow, too. Determining the fully loaded cost of someone in a given position with fifteen years of experience meant reading through one or more full guideline documents.

AI security and privacy

Privacy requirements

The biggest complication: implemented AI solutions had to run on machines the district owns, eliminating the option of using hosted providers such as OpenAI and Anthropic. It also ruled out a hybrid approach where AI runs in a private AWS environment.

Learning the district’s math first

Apple brought our AI team in because of our extensive work in education on the Claris platform. We’re deeply experienced in the client’s technology — The Mac Studio uses unified memory (one pool shared by the processor and graphics chip), which suits running AI models locally.

Our team first sat down with staff members to understand how the district calculates TCO and which questions staff ask most. We used these insights to build fixed logic that returns the same result to these questions every time without fail. The AI solution only presents these results in plain language or a chart, never the number itself. As staff can upload and remove guideline documents themselves, policy answers stay current.

The turning point: one big model became four small ones

We started with the largest open-source model available, at 120 billion parameters, and gave it everything. It worked, but it took about ten seconds just to classify a question. Device records and union contracts were also too different for one model to handle both well.

“The bigger the model, the slower it runs. That is not a flaw. It is physics.”

To address the long wait times, our AI engineers split the work across four smaller models, each sized for one job. After implementation, the time it takes to sort and answer a question dropped to about a second. It may seem small to reduce the wait time from 10 seconds to 1, but it makes a huge difference for the user.

The system is live, and all district data stays on-premises. Local models don’t carry per-token fees, so cost doesn’t exponentially increase as more staff use the assistant. The four-model design further decreased costs from the single-model version it replaced.

AI Assistant
AI + Cloud + FIlemaker

AI + Cloud + FileMaker Experience

Building a full model stack on district hardware takes different skills than calling a cloud API, and it requires deep FileMaker experience. Instead of letting AI guess at a number that drives budget decisions, we built the calculation into the system. We prioritize pivot points in our development process, and when our initial design proved too slow, we changed course quickly and effectively. The separate models framework allows the district to swap in a stronger release later, on its own schedule and without having to rebuild the AI solution.

Under the hood

  • LM Studio and five services in rootless Podman containers all run on one Mac Studio. LM Studio serves the four models with no external API calls.
  • The front end is a React chat app inside a FileMaker web viewer, hosted on FileMaker Server.
  • The Node.js backend routes requests and syncs device data from FileMaker Server into SQLite.
  • A Python FastAPI service runs retrieval-augmented generation (RAG). It stores documents as vectors (numbers that capture meaning) in ChromaDB and retrieves only relevant passages.

Requests try a slash command, then a template match, then the intent classifier. Routine results are cached.

Bring us the AI project that can’t go to the cloud

Many schools and districts still calculate device costs by hand and have policies or union agreements that rule out hosted AI. You can solve both problems on hardware you already own, getting consistent results every time with a custom, local AI solution. Contact our team to talk with an AI consultant about the right AI solution for your privacy requirements and workflows.

Scroll to Top