Skip to content
Menu

Hands-On AI

Trying Jev: From the TypeSafe Playground to an Excuse Court App

We try Jev for the first time, explore its state-and-question interface and build Excuse Court with ChatGPT. The walkthrough shows structured decisions in a small working example, while keeping vendor speed claims separate from what we actually tested.

Watch the first hands-on session

Reindent video thumbnail: Trying Jev: From the TypeSafe Playground to an Excuse Court App

Reindent video · 8 min 23 sec
Watch the full video on YouTube

This September 18 recording is our first use of Jev after receiving access. It moves from the onboarding quiz to the playground and then to a small HTML app built with ChatGPT. It is an exploration of the interface, not a controlled comparison with ChatGPT's speed or accuracy.

How does the state-and-question interface work?

The onboarding makes an important distinction: Jev supplies structured decisions and probabilities, rather than a chat response. The playground separates the information being judged, called the state, from the question and criteria used to judge it.

The first exercise asks whether a food is a sandwich. The definition describes what should count, including a filling enclosed by bread or another structural starch. Changing the food changes the state; changing what qualifies as a sandwich changes the decision criteria.

That separation is the useful idea to carry into another task. It makes the question more explicit than simply asking an assistant for its opinion. It also makes disagreements easier to examine: the input may be ambiguous, the criterion may be unclear or the model may have applied it incorrectly.

What happened when we changed the food?

Diego tries several inputs, including an apple, a burger, an Oreo, sushi, avocado toast and a burrito. The recording shows the outputs responding to those changes. A burger is accepted in the exercise, while an apple and an Oreo are not. The avocado-toast example draws attention to the definition's enclosure requirement.

These examples are playful boundary tests. They do not establish a general measure of factual accuracy. A burrito becoming debatable is a useful reminder that the quality of a classification also depends on the rule being applied. Some disagreements should lead us to improve the specification before blaming or trusting the model.

How did ChatGPT become part of the experiment?

At the quick-start stage, the recording shows TypeSafe's instructions for using Jev with an agent. Diego takes the suggested setup into ChatGPT and uses GPT-6 Astra to help develop a demo. The video then switches to voice chat to choose the idea and request an HTML interface.

The division of work is visible: ChatGPT proposes and builds the application, while Jev evaluates the decisions inside it. This recording is not an installation reference for every account or agent framework. For current setup details, use TypeSafe's documentation rather than assuming that the recorded onboarding screens will stay unchanged.

What does Excuse Court evaluate?

Excuse Court asks three questions about an excuse for missing work: whether it is believable, whether it describes a real emergency and whether it is likely to annoy a boss. The initial example is a cat asleep on a laptop who is supposedly the head of IT.

In the recording, that excuse receives low believability, almost no emergency and moderate predicted annoyance. Diego then tries a broken leg, an alien abduction, being in Hawaii and disliking the job. Those changes make the three judgments diverge, which is the point of the exercise. Something can be believable without being an emergency.

The outputs are model judgments in a toy app. The video does not verify that a person's excuse is true or measure how an actual employer would react. The alien example, in particular, shows why an amusing response should not be mistaken for grounded real-world verification. This is not an employee assessment tool.

Did this prove Jev is 100x faster than ChatGPT?

No. The session feels responsive, but we did not run matched prompts, repeat a timed workload or control the comparison conditions. The speed claim in the video's packaging comes from TypeSafe's launch material for structured-decision tasks. It is not a result measured by this tutorial.

The launch analysis explains the vendor comparisons and their limits. A schema-constrained response also does not guarantee that the selected answer is correct. We did not measure confidence calibration during this first-use session.

What is the practical takeaway?

A small experiment can make an unfamiliar interface understandable. Here, the reusable pattern is to define the input state, ask separate questions and inspect how each answer changes when the input changes. A general-purpose assistant can help build the surrounding interface without turning Jev into a text-generation model.

The next meaningful test would use representative examples with known outcomes and compare the decisions against those outcomes. That work remains ahead of this video. For now, the recording offers a concrete first look at the console and an app that makes its decisions visible.