A Coding Guide to TypeSafe AI Jev: Typed Decisions, Calibrated Confidence, and Speculative Fan-Out with a System One Model
In this tutorial, we work with Jev, TypeSafe AI’s first System One model, which does not generate text at all: we send it a piece of program state and a set of typed questions, and it returns choices, scores, and yes/no probabilities that our code can branch on directly. We install the official Python SDK, make a first call that uses all three question primitives at once, and look at how the shape of the state changes what the model can know. We then recompute the published confidence statistic from the returned probabilities, measure what batching ten questions into one call buys over ten separate calls, and build the patterns the API is designed for: confidence-gated routing, composite scoring with the weights kept in code, typed function calling, and counting done the way the model can actually do it. We close with the production shape: Pydantic response models, an async client fanned out with asyncio, retry policies, typed errors, and a running ledger that prices the whole notebook. Copy CodeCopiedUse a different Browser import os import sys import json import time import asyncio import traceback import subprocess from getpass import getpass RESULTS = {} LEDGER = {“calls”: 0, “input_tokens”: 0, “output_tokens”: 0} USD_PER_MILLION_INPUT_TOKENS = 0.042 # Jev list price; output tokens are free def banner(title): print(“n” + “=” * 78) print(title) print(“=” * 78) def section(name): def wrap(fn): def run(*a, **kw): banner(name) try: out = fn(*a, **kw) RESULTS[name] = out if isinstance(out, str) else “ok” return out except Exception as e: RESULTS[name] = f”SKIPPED / FAILED -> {type(e).__name__}: {e}” print(f”n[!] {name} did not complete: {type(e).__name__}: {e}”) traceback.print_exc(limit=3) return None return run return wrap banner(“0. Install the SDK, load the API key, list the models”) subprocess.run([sys.executable, “-m”, “pip”, “install”, “-q”, “typesafe-sdk==0.7.0”], check=True) import typesafe_sdk from typesafe_sdk import Choice, Noul, Score, TypeSafeClient def load_api_key(): key = os.environ.get(“TYPESAFE_API_KEY”, “”).strip() if not key: try: from google.colab import userdata # Colab: key stored under the Secrets tab key = (userdata.get(“TYPESAFE_API_KEY”) or “”).strip() except Exception: key = “” return key or getpass(“TypeSafe API key (console.typesafe.ai/keys): “).strip() os.environ[“TYPESAFE_API_KEY”] = load_api_key() client = TypeSafeClient() # reads TYPESAFE_API_KEY, defaults to jev-latest print(f” typesafe-sdk {typesafe_sdk.__version__} | Python {sys.version.split()[0]}”) print(” models available to this key:”) for m in client.models.list().models: print(f” {m.name:<14s} released {m.release_date} {m.description}”) def ask(state, questions, **kw): “””One System One call, timed, with its tokens added to the running ledger.””” t0 = time.perf_counter() response = client.system_one(state, questions, **kw) ms = (time.perf_counter() – t0) * 1e3 LEDGER[“calls”] += 1 LEDGER[“input_tokens”] += response.usage.input_tokens or 0 LEDGER[“output_tokens”] += response.usage.output_tokens or 0 return response, ms We install typesafe-sdk, pinned to the version this notebook was written against, and load the API key from the environment, from Colab’s Secrets tab, or from a hidden prompt, so it never appears in the notebook. TypeSafeClient reads TYPESAFE_API_KEY on its own and defaults to the jev-latest alias; listing the models shows which names and pinned versions the key can use. The small ask helper wraps system_one so that every call in the rest of the notebook is timed and its token usage lands in a ledger we total at the end. Copy CodeCopiedUse a different Browser TICKET = { “ticket”: { “subject”: “Duplicate charge”, “messages”: [ {“from”: “customer”, “text”: “I was charged twice for order A-104. This is the second time ” “this year. Please refund the duplicate today.”}, {“from”: “support”, “text”: “We are checking the charges.”}, ], }, “order”: {“id”: “A-104”, “charges”: [{“amount_usd”: 49, “status”: “captured”}, {“amount_usd”: 49, “status”: “captured”}]}, “refund_policy”: “Duplicate charges are eligible for a full refund within 30 days.”, } @section(“1. Three primitives, one call: Choice, Score, Noul”) def three_primitives(): response, ms = ask(TICKET, { “department”: Choice( instructions=”Which team should handle this ticket”, criteria={“billing”: “Payment, refund or subscription issues”, “technical”: “Bugs, outages or integration problems”, “sales”: “Pricing, plans or account upgrades”}, ), “frustration”: Score( instructions=”How frustrated the customer appears in `ticket.messages[0].text`”, criteria=[“Calm, just stating facts”, “Frustrated but civil”, “Very angry, strong language”], ), “refund_requested”: Noul(instructions=”The customer is explicitly asking for a refund”), “policy_supports”: Noul(instructions=”The stated `refund_policy` covers this situation”), }) dept = response.choices[“department”] print(f” department -> {dept.choice!r} confidence {dept.confidence:.3f}”) print(f” probabilities {({k: round(v, 3) for k, v in dept.probabilities.items()})}”) fr = response.scores[“frustration”] print(f” frustration -> score {fr.score:.3f} on 0..{len(fr.legend) – 1} confidence {fr.confidence:.3f}”) for level, text in fr.legend.items(): print(f” {level}: p={fr.probabilities[level]:.3f} {text}”) print(f” refund_requested -> noul {response.nouls[‘refund_requested’].noul:.3f}”) print(f” policy_supports -> noul {response.nouls[‘policy_supports’].noul:.3f}”) print(f”n answered by {response.model} in {ms:.0f} ms ” f”input tokens {response.usage.input_tokens}, output tokens {response.usage.output_tokens}”) return f”{dept.choice}, frustration {fr.score:.2f}, refund {response.nouls[‘refund_requested’].noul:.2f}” three_primitives() A System One request has two parts: state, which is any text, JSON object or array describing the situation, and a dictionary of named questions. Choice selects one label from the criteria we define and returns a probability for every label; Score places the state on an ordered rubric and returns the probability-weighted level, so it can land between two levels; Noul returns a single probability that a statement is true. The question names are ours and never reach the model, which is why the instructions carry the full meaning and can point at nested fields with backticked paths. All four questions are evaluated in one request, in parallel and in isolation from one another, and the response reports the pinned model version that answered and the tokens it billed. Copy CodeCopiedUse a different Browser @section(“2. State is program state: the same question over a string and over named fields”) def state_shapes(): question = {“eligible”: Noul( instructions=”The customer is eligible for a refund under the company’s written policy”, criteria={“true”: “A policy is present and it covers the customer’s situation”, “false”: “No policy is given, or the policy does not cover the situation”}, )} bare = “I was charged twice for order A-104. Please refund the duplicate.” as_list = [m[“text”] for m in TICKET[“ticket”][“messages”]] shapes = [(“string: the message only”, bare), (“array : the conversation”, as_list), (“object: ticket + order + policy”, TICKET)] print(f” {‘state shape’:<34s} {‘noul’:>6s} input tokens ms”) seen = {} for label, state in shapes: response, ms = ask(state, question) seen[label] = response.nouls[“eligible”].noul print(f” {label:<34s} {seen[label]:6.3f} {response.usage.input_tokens:12d} {ms:5.0f}”)




