Skip to content
Get started

How VAKYN works

You send a state and typed questions. VAKYN sends back one answer per question, with a probability. Your code turns that into a decision.

This is the reference version. For the big picture, VAKYN as System 1 next to System 2, see How it works.

Which question type?

Start from what you want to know. Every question is one of three types, and the type decides the shape of the answer.

What are you asking VAKYN?

  • noul

    Yes or no?

    Use for
    A fact or a rule that is either true or false about the state.
    Example
    “Is this product review fake?”
    Returns
    The probability of yes, one number from 0 to 1.
    More about noul
  • choice

    Which one?

    Use for
    Picking one of a few named options: an action, a queue, a category.
    Example
    “Approve, flag or reject this expense report?”
    Returns
    The pick, a probability for each option and a confidence.
    More about choice
  • score

    How much? What level?

    Use for
    Placing the state on an ordered scale: risk, urgency, severity.
    Example
    “How risky is this expense report?”
    Returns
    The expected level, a probability for each level and a confidence.
    More about score
Pick the question type from what you are asking: noul for yes or no, choice for which one, score for how much or what level.

Ask as many questions as you need in one request, and mix the types. The Concepts page covers the fields every question takes.

The flow

One request goes through four steps. The first three happen in VAKYN; the last one is yours.

  1. 01

    State

    “Stopped working after five weeks. 1 star. …”

    Text or JSON. Read once.

  2. 02

    Typed questions

    fakenoultopicchoicenegativescore

    Each one has a type and its options.

  3. 03

    Calibrated answers

    fake
    yes 0.04
    topic
    quality · 0.97
    negative
    3.94 of 4

    Probabilities you can set a threshold on.

  4. 04

    Your decision

    confidence ≥ 0.70?

    Yes: act on it

    No: send to a person

    Your code, your rule.

One request, from the state to the decision: the clear product review from the example below, with real answers from VAKYN-4B Q8_0. Its lowest confidence is 0.95, so the code acts.

1. The state

The state is the thing you want judged: a review, an expense report as JSON, a request an agent wants to make. Pass structure when you have it; field names help the model read it. See State.

2. Typed questions

Each question has a name you choose, a type and its options. The name is only for your code. The type and the options are what VAKYN answers. See Noul, Choice and Score.

3. Calibrated answers

Every answer is a probability distribution: the probability of yes for a noul, a probability per option for a choice, a probability per level for a score. Choice and score answers also carry a confidence. The probabilities are calibrated: an answer given 0.90 is meant to be right about nine times in ten. See Confidence.

4. Your decision

VAKYN answers; your code decides. A typical rule: act when the confidence is at least 0.70, and send everything else to a person. You pick the threshold per question, from how costly a wrong answer is. See Routing by confidence.

Example: a product review

A shop gets reviews all day. For each one it wants three things: is it fake, what is it about, and how negative is it. That is a noul, a choice and a score, asked together about one review. Below are two reviews and the real answers. The clear one is handled by code; the borderline one goes to a moderator.

Questionsfakenoultopicchoicenegativescore
  • Clear case

    Acts automatically
    Stopped working after five weeks. 1 star. I bought the Tallow & Pine CafeDuo 2 in March and used it every morning for two coffees. In week five the steam wand stopped heating, then the pump started rattling and the machine shut itself off mid-shot. I descaled it twice with the brand's own tablets and followed the reset steps in the manual exactly. Nothing helped. For a machine at this price I expected years, not weeks. Do not buy. Marta K., verified purchase
    QuestionAnswerConfidence
    fakeyes 0.040.96
    topicquality0.97
    negative3.94 of 40.95

    Every answer is at least 0.70 sure (lowest: negative 0.95), so your code acts on it.

  • Borderline case

    Goes to a person
    Decent kettle for what I paid. It took ages to show up and the box was a bit crushed, but it boils fine. Would maybe buy again if it goes on sale. J.
    QuestionAnswerConfidence
    fakeyes 0.030.97
    topicquality0.23
    negative1.93 of 40.94

    topic is only 0.23 sure, under 0.70, so a person decides.

Real answers from VAKYN-4B Q8_0. The rule: act when every answer is at least 0.70 sure; a noul's confidence is how far it leans either way.
curl -s "https://api.vakyn.com/v1/systemone" \  -H "Authorization: Bearer $TYPESAFE_API_KEY" \  -H "Content-Type: application/json" \  -o answer.json \  -d @- <<'JSON'{  "model": "jev-latest",  "state": "Stopped working after five weeks. 1 star.\n\nI bought the Tallow & Pine CafeDuo 2 in March and used it every morning for two coffees. In week five the steam wand stopped heating, then the pump started rattling and the machine shut itself off mid-shot. I descaled it twice with the brand's own tablets and followed the reset steps in the manual exactly. Nothing helped. For a machine at this price I expected years, not weeks. Do not buy.\n\nMarta K., verified purchase",  "questions": {    "fake": {      "type": "noul",      "instructions": "Is this review likely fake: paid for, written by the seller, or not based on using the product?"    },    "topic": {      "type": "choice",      "instructions": "What is the review mainly about?",      "criteria": {        "quality": "How the product works, feels or holds up",        "delivery": "Shipping, packaging and arrival",        "price": "Price and value for money",        "service": "Support, returns and refunds",        "other": null      }    },    "negative": {      "type": "score",      "instructions": "How negative is the review?",      "criteria": [        "Positive",        "Mostly positive",        "Mixed",        "Mostly negative",        "Very negative"      ]    }  }}JSON # Act when every answer is at least 0.70 sure; otherwise a person decides.jq -r '.answers as $a  | {fake: ([$a.fake.noul, 1 - $a.fake.noul] | max), topic: $a.topic.confidence, negative: $a.negative.confidence}  | to_entries | min_by(.value) as $weakest  | if $weakest.value >= 0.70    then "act: \(if $a.fake.noul >= 0.5 then "remove" else "publish" end), topic \($a.topic.choice), negative \($a.negative.score) of 4"    else "human: \($weakest.key) is only \($weakest.value * 100 | round / 100) sure" end' answer.json

The borderline case runs the same code with its own state. topic is only 0.23 sure, under 0.70, so it prints a human line and a person decides. Change the threshold to match what a wrong answer costs you.

Run it in the playground, with a threshold slider

Example: an expense report as JSON

Finance checks every expense report before it is paid. The report arrives as JSON, so it goes in as JSON, with the policy next to it. Three questions: approve, flag or reject (a choice), does any line break the policy (a noul), and how risky is paying it as filed (a score). The clear report breaks the policy on every line and goes back to the employee without a person looking at it; the borderline one waits for a person.

Questionsdecisionchoiceover_policynoulriskscore
  • Clear case

    Acts automatically
    { "policy": { "rail": "Second class, booked through the travel desk", "hotel": "Up to EUR 180 a night", "meals": "Up to EUR 60 a day while travelling, no alcohol", "receipts": "A receipt for every line over EUR 25" }, "report": { "employee": "Ines Varga, field engineer", "trip": "Customer site visit in Lyon, 6 October 2026", "lines": [ { "date": "2026-10-06", "item": "Flight, Brussels to Lyon and back, business class, booked privately", "amount_eur": 1240, "receipt": false }, { "date": "2026-10-06", "item": "Hotel, three nights for a one-day visit", "amount_eur": 1260, "receipt": false }, { "date": "2026-10-06", "item": "Dinner, alone, two bottles of wine", "amount_eur": 310, "receipt": false } ], "total_eur": 2810 } }
    QuestionAnswerConfidence
    decisionreject0.82
    over_policyyes 1.001.00
    risk2.97 of 30.97

    Every answer is at least 0.70 sure (lowest: decision 0.82), so your code acts on it.

  • Borderline case

    Goes to a person
    { "policy": { "rail": "Second class, booked through the travel desk", "hotel": "Up to EUR 180 a night", "meals": "Up to EUR 60 a day while travelling, no alcohol", "receipts": "A receipt for every line over EUR 25" }, "report": { "employee": "Pieter Claes, account manager", "trip": "Trade fair in Milan, 13 to 14 October 2026", "lines": [ { "date": "2026-10-13", "item": "Rail, Brussels to Milan, second class, booked by the travel desk", "amount_eur": 164, "receipt": true }, { "date": "2026-10-13", "item": "Hotel, one night, city tax EUR 9 included", "amount_eur": 189, "receipt": true }, { "date": "2026-10-13", "item": "Dinner with two people from a prospect", "amount_eur": 118, "receipt": true, "note": "Talked about a pilot" }, { "date": "2026-10-14", "item": "Taxi to the fair", "amount_eur": 27, "receipt": false, "note": "Driver's card reader was broken" } ], "total_eur": 498 } }
    QuestionAnswerConfidence
    decisionreject0.65
    over_policyyes 0.940.94
    risk2.52 of 30.52

    risk is only 0.52 sure, under 0.70, so a person decides.

Real answers from VAKYN-4B Q8_0. The rule: act when every answer is at least 0.70 sure; a noul's confidence is how far it leans either way.
curl -s "https://api.vakyn.com/v1/systemone" \  -H "Authorization: Bearer $TYPESAFE_API_KEY" \  -H "Content-Type: application/json" \  -o answer.json \  -d @- <<'JSON'{  "model": "jev-latest",  "state": {    "policy": {      "rail": "Second class, booked through the travel desk",      "hotel": "Up to EUR 180 a night",      "meals": "Up to EUR 60 a day while travelling, no alcohol",      "receipts": "A receipt for every line over EUR 25"    },    "report": {      "employee": "Ines Varga, field engineer",      "trip": "Customer site visit in Lyon, 6 October 2026",      "lines": [        {          "date": "2026-10-06",          "item": "Flight, Brussels to Lyon and back, business class, booked privately",          "amount_eur": 1240,          "receipt": false        },        {          "date": "2026-10-06",          "item": "Hotel, three nights for a one-day visit",          "amount_eur": 1260,          "receipt": false        },        {          "date": "2026-10-06",          "item": "Dinner, alone, two bottles of wine",          "amount_eur": 310,          "receipt": false        }      ],      "total_eur": 2810    }  },  "questions": {    "decision": {      "type": "choice",      "instructions": "What should finance do with this expense report?",      "criteria": {        "approve": "Pay it as filed",        "flag": "Hold it until a person checks it",        "reject": "Send it back to the employee to fix"      }    },    "over_policy": {      "type": "noul",      "instructions": "Does any line in the report break the expense policy?"    },    "risk": {      "type": "score",      "instructions": "How risky is it to pay this report as filed?",      "criteria": [        "No risk",        "Low risk",        "Some risk",        "High risk"      ]    }  }}JSON # Act when every answer is at least 0.70 sure; otherwise a person decides.jq -r '.answers as $a  | $a | map_values(if .type == "noul" then [.noul, 1 - .noul] | max else .confidence end)  | to_entries | min_by(.value) as $weakest  | if $weakest.value >= 0.70    then "act: \($a.decision.choice), over policy \(if $a.over_policy.noul >= 0.5 then "yes" else "no" end), risk \($a.risk.score) of 3"    else "human: \($weakest.key) is only \($weakest.value * 100 | round / 100) sure" end' answer.json

The borderline case runs the same code with its own state. risk is only 0.52 sure, under 0.70, so it prints a human line and a person decides. Change the threshold to match what a wrong answer costs you.

Run it in the playground, with a threshold slider

Example: a guardrail for an agent

An agent that handles returns wants to refund EUR 900 through the payments API. Before the call goes out, VAKYN reads the call, the agent's job and what the shop knows, and answers four questions: allow, ask a person or block (a choice), is it hard to undo (a noul), how much damage could a mistake do (a score), and is it part of the agent's job (a noul). The clear call goes through; the borderline one waits for a person to approve it.

Questionsactionchoiceirreversiblenouldamagescorewithin_jobnoul
  • Clear case

    Acts automatically
    { "agent": { "name": "Returns assistant at Orla Cycles", "job": "Handles returns. May refund returned items to the original payment method, up to EUR 1,000 per order, once the warehouse has checked the return in." }, "call": { "api": "POST /payments/refunds", "body": { "order": "OC-58213", "amount": 900, "currency": "EUR", "to": "original payment method", "reason": "Returned, unused" } }, "context": { "order": "OC-58213, one e-bike battery, paid EUR 900.00 by card on 2026-09-28", "return": "Checked in by the warehouse on 2026-10-09: sealed, unused, serial number matches the order", "earlier_refunds_on_order": 0, "customer": "Customer since 2022, 11 orders, no earlier refunds, same card on file since 2022" } }
    QuestionAnswerConfidence
    actionallow1.00
    irreversibleyes 0.140.86
    damage0.21 of 30.79
    within_jobyes 0.990.99

    Every answer is at least 0.70 sure (lowest: damage 0.79), so your code acts on it.

  • Borderline case

    Goes to a person
    { "agent": { "name": "Returns assistant at Orla Cycles", "job": "Handles returns. May refund returned items to the original payment method, up to EUR 1,000 per order, once the warehouse has checked the return in." }, "call": { "api": "POST /payments/refunds", "body": { "order": "OC-60477", "amount": 900, "currency": "EUR", "to": "original payment method", "reason": "Parcel never arrived" } }, "context": { "order": "OC-60477, one e-bike battery, paid EUR 900.00 by card on 2026-10-02", "return": "None: the customer says the parcel never arrived", "carrier": "Tracking says delivered to a parcel locker on 2026-10-05", "earlier_refunds_on_order": 0, "customer": "First order, account opened 2026-10-01" } }
    QuestionAnswerConfidence
    actionblock0.93
    irreversibleyes 0.290.71
    damage0.77 of 30.23
    within_jobyes 0.420.58

    damage is only 0.23 sure, under 0.70, so a person decides.

Real answers from VAKYN-4B Q8_0. The rule: act when every answer is at least 0.70 sure; a noul's confidence is how far it leans either way.
curl -s "https://api.vakyn.com/v1/systemone" \  -H "Authorization: Bearer $TYPESAFE_API_KEY" \  -H "Content-Type: application/json" \  -o answer.json \  -d @- <<'JSON'{  "model": "jev-latest",  "state": {    "agent": {      "name": "Returns assistant at Orla Cycles",      "job": "Handles returns. May refund returned items to the original payment method, up to EUR 1,000 per order, once the warehouse has checked the return in."    },    "call": {      "api": "POST /payments/refunds",      "body": {        "order": "OC-58213",        "amount": 900,        "currency": "EUR",        "to": "original payment method",        "reason": "Returned, unused"      }    },    "context": {      "order": "OC-58213, one e-bike battery, paid EUR 900.00 by card on 2026-09-28",      "return": "Checked in by the warehouse on 2026-10-09: sealed, unused, serial number matches the order",      "earlier_refunds_on_order": 0,      "customer": "Customer since 2022, 11 orders, no earlier refunds, same card on file since 2022"    }  },  "questions": {    "action": {      "type": "choice",      "instructions": "What should the guardrail do with this API call?",      "criteria": {        "allow": "Let the agent make the call now",        "ask_human": "Hold the call until a person approves it",        "block": "Stop the call and tell the agent why"      }    },    "irreversible": {      "type": "noul",      "instructions": "Would this call be hard or impossible to undo once made?"    },    "damage": {      "type": "score",      "instructions": "If this call is a mistake, how hard is the damage to put right?",      "criteria": [        "Easy",        "Takes some work",        "Hard",        "Impossible"      ]    },    "within_job": {      "type": "noul",      "instructions": "Is this call part of the agent's job as described?"    }  }}JSON # Act when every answer is at least 0.70 sure; otherwise a person approves the call.jq -r '.answers as $a  | $a | map_values(if .type == "noul" then [.noul, 1 - .noul] | max else .confidence end)  | to_entries | min_by(.value) as $weakest  | if $weakest.value >= 0.70    then "act: \($a.action.choice) (damage \($a.damage.score) of 3)"    else "human: \($weakest.key) is only \($weakest.value * 100 | round / 100) sure" end' answer.json

The borderline case runs the same code with its own state. damage is only 0.23 sure, under 0.70, so it prints a human line and a person decides. Change the threshold to match what a wrong answer costs you.

Run it in the playground, with a threshold slider

Example: eight job applications in one request

One request can ask many questions about one state. Here the state holds eight applications for one role, and the request asks ten questions: a priority score for each application, which one to interview first (a choice) and whether any of them shows a red flag (a noul). The state is read once, so the ten questions cost little more than one. The code builds the per-application questions from the data.

Questionspriority_1scorepriority_2scorepriority_3scorepriority_4scorepriority_5scorepriority_6scorepriority_7scorepriority_8scorebestchoicered_flagsnoul
  • Clear case

    Acts automatically
    { "role": { "title": "Night shift lead, distribution center in Rotterdam", "must_have": [ "Two years or more leading a warehouse team", "A valid forklift certificate", "Can work nights, Sunday to Thursday" ], "nice_to_have": [ "Dutch and English", "Has used a warehouse management system" ] }, "applications": [ { "application": 1, "name": "Joost Verhagen", "summary": "Four years as night shift lead at a grocery distribution center, team of 14. Forklift certificate valid to 2028. Wants to stay on nights. Dutch and English. Daily work in a warehouse management system." }, { "application": 2, "name": "Mila Petrova", "summary": "Recent graduate in logistics, no work experience in a warehouse yet. No forklift certificate. Prefers day shifts." }, { "application": 3, "name": "Sam O'Brien", "summary": "Barista for three years. No warehouse or forklift experience. Available weekends only." }, { "application": 4, "name": "Fatima Zahra El Idrissi", "summary": "Office manager for six years. No warehouse experience, no forklift certificate. Looking for a day job close to home." }, { "application": 5, "name": "Kees de Wit", "summary": "Retired in 2025 after thirty years as a truck driver. No team lead experience. Wants two days a week." }, { "application": 6, "name": "Lucas Moreau", "summary": "Software developer looking for a career change into tech sales. Never worked in a warehouse." }, { "application": 7, "name": "Hanna Kowalski", "summary": "Student, wants a summer job. No experience and no forklift certificate. Available in July and August only." }, { "application": 8, "name": "Ravi Menon", "summary": "Chef for ten years in restaurants. No warehouse experience. Cannot work nights." } ] }
    QuestionAnswerConfidence
    priority_11.98 of 20.97
    priority_20.00 of 21.00
    priority_30.00 of 21.00
    priority_40.00 of 21.00
    priority_50.01 of 20.99
    priority_60.00 of 21.00
    priority_70.00 of 21.00
    priority_80.00 of 21.00
    bestapplication_10.96
    red_flagsyes 0.090.91

    Every answer is at least 0.70 sure (lowest: red_flags 0.91), so your code acts on it.

  • Borderline case

    Goes to a person
    { "role": { "title": "Night shift lead, distribution center in Rotterdam", "must_have": [ "Two years or more leading a warehouse team", "A valid forklift certificate", "Can work nights, Sunday to Thursday" ], "nice_to_have": [ "Dutch and English", "Has used a warehouse management system" ] }, "applications": [ { "application": 1, "name": "Anouk Jansen", "summary": "Three years as shift lead in a parcel hub, team of 9, mostly evenings. Forklift certificate valid to 2027. Open to nights. Dutch and English." }, { "application": 2, "name": "Daniel Mensah", "summary": "Five years as warehouse team lead on day shifts, team of 20. Forklift certificate expired in 2025, says he will renew it. Can work nights. English only." }, { "application": 3, "name": "Eva Lindgren", "summary": "Two years leading a night team of 6 in a cold store. Forklift certificate valid. Uses a warehouse management system daily. English and some Dutch." }, { "application": 4, "name": "Tom Bakker", "summary": "Forklift driver for seven years, no team lead role. Certificate valid. Night shifts now. Dutch." }, { "application": 5, "name": "Yusuf Demir", "summary": "Says he led a team of 30 for eight years; his dates show four years in total at two companies. Forklift certificate attached." }, { "application": 6, "name": "Sofia Russo", "summary": "Two years as assistant shift lead at a fashion warehouse. Forklift certificate valid. Prefers nights. English and Italian." }, { "application": 7, "name": "Bram Visser", "summary": "Nine years as a warehouse supervisor, then three years out of work for family reasons. Certificate valid to 2026. Can work nights." }, { "application": 8, "name": "Lena Hofmann", "summary": "Store manager in retail for five years, team of 12. No forklift certificate. Can work nights." } ] }
    QuestionAnswerConfidence
    priority_11.85 of 20.77
    priority_20.29 of 20.56
    priority_31.81 of 20.71
    priority_40.05 of 20.93
    priority_50.77 of 20.27
    priority_60.12 of 20.82
    priority_71.00 of 20.49
    priority_80.01 of 20.99
    bestapplication_30.49
    red_flagsyes 0.930.93

    priority_5 is only 0.27 sure, under 0.70, so a person decides.

Real answers from VAKYN-4B Q8_0. The rule: act when every answer is at least 0.70 sure; a noul's confidence is how far it leans either way.
curl -s "https://api.vakyn.com/v1/systemone" \  -H "Authorization: Bearer $TYPESAFE_API_KEY" \  -H "Content-Type: application/json" \  -o answer.json \  -d @- <<'JSON'{  "model": "jev-latest",  "state": {    "role": {      "title": "Night shift lead, distribution center in Rotterdam",      "must_have": [        "Two years or more leading a warehouse team",        "A valid forklift certificate",        "Can work nights, Sunday to Thursday"      ],      "nice_to_have": [        "Dutch and English",        "Has used a warehouse management system"      ]    },    "applications": [      {        "application": 1,        "name": "Joost Verhagen",        "summary": "Four years as night shift lead at a grocery distribution center, team of 14. Forklift certificate valid to 2028. Wants to stay on nights. Dutch and English. Daily work in a warehouse management system."      },      {        "application": 2,        "name": "Mila Petrova",        "summary": "Recent graduate in logistics, no work experience in a warehouse yet. No forklift certificate. Prefers day shifts."      },      {        "application": 3,        "name": "Sam O'Brien",        "summary": "Barista for three years. No warehouse or forklift experience. Available weekends only."      },      {        "application": 4,        "name": "Fatima Zahra El Idrissi",        "summary": "Office manager for six years. No warehouse experience, no forklift certificate. Looking for a day job close to home."      },      {        "application": 5,        "name": "Kees de Wit",        "summary": "Retired in 2025 after thirty years as a truck driver. No team lead experience. Wants two days a week."      },      {        "application": 6,        "name": "Lucas Moreau",        "summary": "Software developer looking for a career change into tech sales. Never worked in a warehouse."      },      {        "application": 7,        "name": "Hanna Kowalski",        "summary": "Student, wants a summer job. No experience and no forklift certificate. Available in July and August only."      },      {        "application": 8,        "name": "Ravi Menon",        "summary": "Chef for ten years in restaurants. No warehouse experience. Cannot work nights."      }    ]  },  "questions": {    "priority_1": {      "type": "score",      "instructions": "How high should application 1 go on the interview list?",      "criteria": [        "Do not interview",        "Interview if needed",        "Interview first"      ]    },    "priority_2": {      "type": "score",      "instructions": "How high should application 2 go on the interview list?",      "criteria": [        "Do not interview",        "Interview if needed",        "Interview first"      ]    },    "priority_3": {      "type": "score",      "instructions": "How high should application 3 go on the interview list?",      "criteria": [        "Do not interview",        "Interview if needed",        "Interview first"      ]    },    "priority_4": {      "type": "score",      "instructions": "How high should application 4 go on the interview list?",      "criteria": [        "Do not interview",        "Interview if needed",        "Interview first"      ]    },    "priority_5": {      "type": "score",      "instructions": "How high should application 5 go on the interview list?",      "criteria": [        "Do not interview",        "Interview if needed",        "Interview first"      ]    },    "priority_6": {      "type": "score",      "instructions": "How high should application 6 go on the interview list?",      "criteria": [        "Do not interview",        "Interview if needed",        "Interview first"      ]    },    "priority_7": {      "type": "score",      "instructions": "How high should application 7 go on the interview list?",      "criteria": [        "Do not interview",        "Interview if needed",        "Interview first"      ]    },    "priority_8": {      "type": "score",      "instructions": "How high should application 8 go on the interview list?",      "criteria": [        "Do not interview",        "Interview if needed",        "Interview first"      ]    },    "best": {      "type": "choice",      "instructions": "Which application should we interview first?",      "criteria": {        "application_1": null,        "application_2": null,        "application_3": null,        "application_4": null,        "application_5": null,        "application_6": null,        "application_7": null,        "application_8": null      }    },    "red_flags": {      "type": "noul",      "instructions": "Does any application show a red flag, such as claims that contradict each other or a certificate that looks made up?"    }  }}JSON # Act when every answer is at least 0.70 sure; otherwise a person decides.jq -r '.answers as $a  | $a | map_values(if .type == "noul" then [.noul, 1 - .noul] | max else .confidence end)  | to_entries | min_by(.value) as $weakest  | if $weakest.value >= 0.70    then "act: interview \($a.best.choice) first, red flags \(if $a.red_flags.noul >= 0.5 then "yes" else "no" end), "      + "priorities \([range(1; 9) as $i | $a["priority_\($i)"].score | round] | map(tostring) | join(" "))"    else "human: \($weakest.key) is only \($weakest.value * 100 | round / 100) sure" end' answer.json

The borderline case runs the same code with its own state. priority_5 is only 0.27 sure, under 0.70, so it prints a human line and a person decides. Change the threshold to match what a wrong answer costs you.

Run it in the playground, with a threshold slider

Under the hood

What happens inside one request, in plain words.

State

Read once

The model reads the state one time and keeps that reading in a cache.

cache

  • fake noulHead: scores yes and noCalibrated: for noul, 2 optionsyes 0.04
  • topic choiceHead: scores 5 topicsCalibrated: for choice, 5 optionsquality · 0.97
  • negative scoreHead: scores 5 levelsCalibrated: for score, 5 levels3.94 of 4

Each question continues from the cached state, so only its own words are read.

Runs onVAKYN MAX, hosted by usOr your own GPU or CPU, with open UQFF weights (Apache-2.0)Self-hosted, nothing leaves the machine
Inside one request: the state is read once and cached; each question branches from the cache; a small pointer head scores each option from the model's hidden states; each answer is calibrated for its question type and number of options; it runs on VAKYN MAX or, with the open weights, on your own hardware.
  1. The state is read once. VAKYN reads it a single time and keeps that reading in a cache.
  2. Each question branches from the cache. Only the question and its options are read on top of the state. That is why ten questions about one document cost little more than one.
  3. A small pointer head scores the options. It looks at the model's hidden states (its internal reading) at each option and at the point of decision, and gives each option a score. No text is generated, so there is nothing to parse.
  4. Each answer is calibrated. The scores become probabilities, adjusted for the question type and the number of options, so a confidence means the same across your questions.
  5. It runs where you choose. On VAKYN MAX, our hosted API, or on your own GPU or CPU with open-server. The weights are open (UQFF, Apache-2.0); self-hosted, the state and the answers never leave the machine. See Self-hosting.

Loading the docs…