Four apps where the model has to show its work

I built four tools for chores that eat time: interview practice, meeting notes, bill splitting and release notes. Each one makes the model cite its source or show its sums, and plain code checks the answer before you see it.

I wanted a set of projects that people would actually use, not demos of a model being clever. So I picked four chores that eat time: practicing for an interview, writing up a meeting, splitting a restaurant bill, and writing release notes. Each one is live: Interview Coach, Minutes, SplitSnap and ShipNotes. All four run on the OpenAI API, and the code is on my GitHub.

The rule I gave myself was that the model has to show its work, and plain code has to check it. If an answer cites a line, the line has to exist. If a receipt has a total, the items have to add up to it. A model can be wrong in a confident voice, so every app has one check it cannot talk its way past.

Interview Coach

The sample report: each answer scored for structure, clarity and depth, with the quotes that back each score.

Paste a job description, pick a level, and talk to an interviewer for six minutes in the browser. The voice runs over the OpenAI Realtime API on WebRTC, with a live transcript and an orb that reacts to both voices. When the call ends, a second model reads the transcript and returns a structured report: every question, a score for structure, clarity and depth, evidence quotes, and one thing to improve.

The check is on the quotes. Each one is matched against what the candidate actually said, and quotes that do not match are hidden. A score above 3 with no verified quote is capped at 3, and the report says so. In one test 17 of 18 quotes matched. The one that failed was a paraphrase: the model wrote "in the 75th percentile" where the candidate had said "at the". That is exactly the kind of small invention the check exists for.

Voice is also the expensive part, so the browser never gets a key. The server opens the realtime call itself and hangs it up when time runs out. I tested that by keeping a call open past the timer on purpose. The server ended it at about 377 seconds.

Minutes

A product standup becomes action items with owners and dates, and clicking one jumps to the line it came from.

Record or upload a meeting and get a summary, decisions, action items with owners and due dates, open questions, risks and a follow up email. Transcription uses a diarizing model, so each line has a speaker. You can export to markdown or to a calendar file for the action items.

Every item has to cite the transcript lines it came from, plus a short quote. Unknown line ids are dropped, items with no valid citation are removed, and an item whose quote does not match its lines gets a "check source" badge. Clicking an item highlights the line and seeks the audio to it. On the sample standup all 12 items passed.

Two things I learned. Similar voices got merged into one speaker until I sent short reference clips for each person. And transcription is billed by audio length, not file size, so the server reads the duration from the file's header before calling OpenAI and keeps a daily budget of minutes.

SplitSnap

A dinner receipt read, checked, assigned to four people and settled to the exact cent.

Take a photo of a receipt, check what was read, tap who had what, and send everyone their share. A vision model reads the receipt into typed fields. I told it to copy the printed numbers exactly and not fix the arithmetic, because the app checks the arithmetic itself: quantity times price for each line, lines against the subtotal, and subtotal plus tax and tip against the total.

One sample is a handwritten bill with a real mistake, 2 x 3.20 written as 7.40. The model copied it as printed and the app flagged the line with a one tap fix. All money is integer cents, shared dishes split evenly, and tax and tip are spread with a largest remainder rule, so the shares always add up to the total. A test runs 3,000 random receipts to prove it. On the sample dinner the four shares come to exactly $342.81. The split lives in the share link, so nothing is stored on a server.

ShipNotes

Release notes for a real open source release, grouped and linked, with a tone switch for users or developers.

Paste a public GitHub repo and pick two tags. The server pulls the commits and merged pull requests in that range, and the model groups them into breaking changes, features, fixes and so on, writing each one twice: once for users and once for developers. You get a changelog page, a markdown export and a draft GitHub release body.

Every bullet has to cite the pull requests or commits it describes. Citations outside the fetched range are dropped, a bullet left with no source is removed, and any change the model skipped is added back from its own title with a small label. On the ky sample, 13 changes came back with 13 citations and none invented. A run costs well under two cents. One dull bug was worth catching too: sorting tags by date made the default range pick a backport of an old version, so tags are now sorted by version number.

What these have in common

A security review found real problems in the first versions, and fixing them taught me as much as building the features. The rate limiter could be reset by flooding it with new keys, request bodies were read in full before their size was checked, and the voice session cap was only enforced by the browser. All of those are fixed now, and each fix has a test.

Try them: Interview Coach, Minutes, SplitSnap and ShipNotes. The code is at interview-coach, minutes, splitsnap and shipnotes.