Five apps on a model that only answers in numbers

I built five more apps on TypeSafe's Jev in a day: a tone checker, a headline ranker, a contract reviewer, a fallacy finder and a pitch panel. Every one needed its thresholds moved after looking at real numbers, and that turned out to be the job.

After JobFit I wanted to know whether TypeSafe's Jev was a one trick model or a real building block. Jev never writes a word. You give it some state and a set of typed questions, and it gives you back probabilities and scores. So I built five more apps on it in a day, each one a different kind of judgement: ToneRadar for how a message will land, Headline Arena for which headline wins the click, FinePrint for risky contract clauses, fallacy finder for weak arguments, and PitchPanel for startup pitches. All five are open source on my GitHub.

The same shape five times

Every app follows one pattern. Split the input into units a person would judge one at a time, like sentences, clauses or headlines. Send each unit to Jev with the context it needs and a handful of questions. Then do everything a reader sees in plain code: the scoring formula, the ranking, the chart and every sentence of explanation. The model decides how likely something is. My code decides what that means and how to say it.

That split turned out to be the useful part. Because no sentence on any page comes from the model, nothing can be invented. A clause is flagged because a probability crossed a line I chose, and the explanation under it is one I wrote and can test. Each app has unit tests for its scoring logic, which is hard to say about a prompt.

ToneRadar

A snippy Slack message and its rewrite on one radar, with the sentences to fix first.

Paste a message you are about to send and pick who it is for. Jev scores the whole message on eight tone axes such as warmth, clarity and passive aggression, and asks a few yes or no questions about every sentence, like whether it could be read as angry or whether it is a clear ask. The radar, the heat marked text and the list of sentences to rewrite all come from those numbers. In the sample, the original Slack draft lands at 32 and the rewrite at about 80.

The first surprise was greetings. A line like "Hi Dana," inside a tense email came back as high risk, because Jev was judging the sentence in the mood of the message around it. The fix was one more boolean per sentence, whether it carries real content, which lets greetings and sign offs count for less.

Headline Arena

Five engineering blog titles go in, a champion and a full standings table come out.

Put two to eight headlines in the ring and they get ranked like a sports table. Each one is scored on curiosity, clarity, specificity, emotional pull and credibility, plus three booleans: is it clickbait, would this audience click, and does it promise something they want. The power rating mixes those and takes off up to 30 points for clickbait.

My first clickbait question was too loose. An honest benefit headline came back at 0.80 clickbait, so it lost to worse ones. I rewrote the question with explicit criteria for what counts as clickbait and moved the penalty threshold from 35 percent to 50. After that, "You won't BELIEVE what this designer charged" scored 0.94 clickbait and finished last, and the specific how to headline won.

FinePrint

A streaming service's terms, marked up clause by clause with a risk gauge and the three to read first.

Paste a lease, an offer letter or a set of terms, or upload a PDF. Each clause gets seven risk booleans, such as auto renewal, arbitration and data sharing, a score for how aggressive it is next to a normal contract, and a boolean for whether a typical person needs to notice it. The page reads like a redlined document, with tinted clauses, margin notes and a gauge for the whole contract. It says clearly that it is not legal advice.

Two calibration notes. Jev rates arbitration clauses as standard for terms of service, which is accurate and also exactly what a reader wants flagged, so a category hit now carries more weight than how unusual the wording is. And the "needs to notice" question came back between 0.8 and 0.95 for almost every clause. A signal that is always high cannot sort risk, so it only boosts the ranking of the read first list.

fallacy finder

A two speaker debate about bike lanes, with the ad hominem and strawman lines underlined.

Paste an argument, an op ed or a thread. Lines that start with A: and B: become a two speaker debate with a scoreboard. Every sentence is checked for eight common fallacies, whether it makes a factual claim, and whether it backs that claim up. Click a flagged line and you get a card with a definition and a classic example, both written by hand.

I tested it on an op ed with one fallacy planted per line. It caught seven of eight with probabilities between 0.84 and 0.96. The miss was whataboutism, which topped out at 0.29. Jev also likes to fire strawman next to other fallacies, so a line shows its strongest match and hides the rest behind a small count.

PitchPanel

A grid software pitch judged by five investor types, each with a score, a verdict and a reason.

Paste a pitch and five investor archetypes judge it: an operator, a market hawk, a skeptic, a product person and a numbers person. Each one is the same ten criteria with a different lens in the state, so the skeptic weighs defensibility and the market hawk weighs timing. You get in, maybe or pass from each, a heatmap of criteria against judges, and advice for the weakest areas. Pitches are saved in the browser so you can compare versions.

The red flag question returned about 0.45 to 0.5 even for the strongest sample pitch. Subtracting it directly punished noise, so the penalty only applies above 0.5. The strong sample now scores in the 70s with four of five judges asking for a second meeting, and the rough one scores 11.

What I would tell someone trying this

Keeping a public demo cheap

These run on my own credits, so each app allows five runs per visitor per hour and checks the input before it counts against that limit. A security review caught that the first version read the visitor's IP from the leftmost x-forwarded-for entry, which a caller can set to anything. All five now take it from the header the proxy sets. FinePrint also checks a PDF's size before reading the upload, not after.

Try them: ToneRadar, Headline Arena, FinePrint, fallacy finder and PitchPanel. The code for each is on GitHub at toneradar, headline-arena, fineprint, fallacy-finder and pitchpanel.