JEVANY / DOCUMENTATION

Get started with JevAny

JevAny is open infra for System 1 decision model training and deployment, covering data preparation, model adaptation and evaluation. Use a released model or train on your own data to route support tickets, select tools, or choose a robot's next action. One API takes the state, question and candidate options, then directly returns a choice and its probabilities.

JevAny infra for System 1 decision model training, deployment and application integration

🎮 Results and Demos

Explore interactive benchmark results

Full benchmark results and evaluation details.

The following 30 examples are archived replays from an earlier compatible JevAny checkpoint. The current default release is JevAny-Qwen3.8-27B. Explore the cases, or run a model locally to try your own inputs and see its choices and probabilities.

Explore all 30 application replays →

⚡ Jev inside LLM agent loops

Jev chooses among valid, reversible actions proposed by the LLM, which handles planning, recovery, and completion. The animations compare LLM only (left) with LLM + Jev (right) at equal reward. Steps are illustrated; accelerated playback preserves each pair's measured completion-time ratio. Click an animation to enlarge it.

Golden rules

1. WebShop
Jev selects the requested color and size from LLM-generated menus before the LLM buys the product. Actions fall from 9 to 5, LLM calls from 9 to 4, tokens from 38,852 to 14,256, and time from 18.54 to 7.83 seconds.

2. FrozenLake
After one LLM plan, Jev checks each new state and chooses among four directions. Both runs reach the goal in 4 moves; LLM decision calls fall from 4 to 1, tokens from 2,338 to 663, and time from 19.7 to 16.7 seconds.

3. Terminal-Bench
For sqlite-db-truncate, Jev selects raw-page inspection from three commands, then the LLM recovers and verifies ten rows. Tool commands fall from 13 to 7, LLM calls from 15 to 8, tokens from 202,050 to 121,293, and time from 187.9 to 144.7 seconds.

Broader paired evaluations show that gains vary by task. The full results, delegation protocol, and technical report describe where Jev helps and when to return control to the LLM.

Task Success Efficiency
FrozenLake (GPT-5.6-sol, 10 pairs) 100% → 100% LLM calls −64.4%, tokens −63.1%, time −37.6%
WebShop (LLM-generated menus, 3 pairs) 67% → 100% LLM calls −21.4%, tokens −14.3%, time −15.0%
WebArena (6 pairs) 50% → 50% LLM calls +5.6%, tokens +28.2%, time −0.4%
Terminal-Bench (6 pairs) 1/6 → 3/6 LLM calls −9.0%

📑 Table of Contents

⚡ 1. Quickstart

Use Python 3.12 or newer. Clone the repository and install the lightweight package:

git clone https://github.com/SimpleJev/JevAny.git
cd JevAny
python3.12 -m venv .venv
source .venv/bin/activate
python -m pip install -e .

Keep this environment active and work from the repository root. Start with a local demo, then train on your own data or use the API.

💻 1.1 Run locally

Choose a model that fits your computer:

Model Hardware Start here
Qwen 0.8B starter CPU · 16 GB RAM recommended Train the small adapter on the bundled tickets
JevAny-Qwen 4B CUDA · ~8 GB for BF16 base weights, plus runtime memory Load the released model
JevAny-Qwen 27B CUDA · ~54 GB for BF16 base weights, plus runtime memory Choose the larger checkpoint

The local model guide covers preparation and loading. Released models download on first use and reuse the local cache. With the model server running, open a second terminal in the same checkout:

source .venv/bin/activate
jevany demo --base-url http://127.0.0.1:8008 --text-only

Open http://127.0.0.1:8090, choose Test and connect, then edit Try your own decision and press Ask the model. Change the state or options to see how its decision changes. Games, robotics and replays are available in the same playground.

🛠️ 1.2 JevAny Training

Train your own System 1 model on the same state and questions you send at inference, with a label for each question. Start with the bundled synthetic support tickets, then train on your own labelled data. The starter recipe uses Qwen3.5-0.8B on CUDA with BF16 and writes runs/my-jev:

python -m pip install -e '.[train]'
jevany data init --out data/starter
jevany data validate data/starter/train.jsonl
jevany train --config recipes/sft.toml --dry-run
jevany train --config recipes/sft.toml

After training, try the checkpoint on the included ticket request:

jevany decide examples/request.json --checkpoint runs/my-jev

Pass --data to train on your own JSONL data, or use recipes/finetune.toml to adapt the released 27B model. See the training guide for CPU settings, multimodal data and standard torchrun launches. For image/video training or fine-tuning the released 27B model, install .[train,multimodal].

After SFT, you can continue with experimental RLCR, which rewards correctness and probability calibration:

jevany train --config recipes/rlcr.toml

🚀 1.3 JevAny Deployment

Install the serving dependencies and start the released Qwen 4B model on a CUDA GPU. See the hardware and loading guide for memory requirements.

python -m pip install -e '.[serve,multimodal]'
jevany serve --checkpoint SimpleJev/JevAny-Qwen3.5-4B-LoRA \
  --device cuda --dtype bf16 --port 8008

The default path favors reproducibility. CUDA deployments can opt into BF16 LoRA merging, SDPA and torch.compile; the useful settings differ between 4B and 27B. See the inference acceleration guide for commands, H200 measurements and accuracy caveats.

To serve your training output, replace the checkpoint ID with runs/my-jev. Keep the server running. In a Python session using the same environment, send a ticket and the departments that can handle it:

from jevany import Choice, JevClient

jev = JevClient("http://127.0.0.1:8008")
result = jev.system_one(
    state={"ticket": "I was charged twice. Please help."},
    questions={
        "department": Choice(
            instructions="Which team should handle this?",
            criteria={"billing": "Payment problems", "shipping": "Delivery problems"},
        ),
    },
)
answer = result["answers"]["department"]
print("Selected team:", answer["choice"])
print("Probabilities:", answer["probabilities"])

choice is one of the department names; probabilities maps each name to its probability. Your application can use these fields to route the ticket or ask for review when the decision is uncertain. Use Noul for yes/no questions, such as whether a ticket needs urgent review, and Score for ordered levels, such as low, normal and high priority. See the API reference for all three question types.

For in-process inference, load a model in Python and use the same interface. For image and video inputs, follow the media setup.

🤗 2. Pretrained Models

For a first local run, choose a model and hardware in Run locally.

Model Readout Intended use
 JevAny-Gemma-4B Pointer Compact Gemma release
 JevAny-Qwen3.5-4B Pointer Compact, flexible choice count
 JevAny-Qwen3.5-4B-Direct-Token Direct-token Best released 4B JevBench accuracy
 JevAny-Qwen3.8-27B Pointer Default; highest released accuracy
 JevAny-Muse-Glimmer-30B Pointer Muse Glimmer alternative

These LoRA adapters were trained with SFT on 1,772,725 text records containing 2,180,242 labelled decisions; see training compute and experiments for the setup. Full-parameter SFT and further post-training improvements are planned.

The corresponding base model is loaded separately and its license and access terms apply. Allow roughly twice the base parameter count in bytes for BF16 weights, plus runtime memory. See the hardware and loading guide.

Pointer and direct-token models share the same API. Pointer supports up to 4,096 options within the context limit; direct-token supports up to 255. See readout choices for training and accuracy tradeoffs.

📊 3. Benchmark Results

JevAny-Qwen3.8-27B leads both benchmarks and has the lowest NLL and Brier. Among 4B releases, direct-token leads on JevBench; pointer leads on Transfer.

Explore interactive benchmark results

| Model | Transfer ↑ | JevBench ↑ | NLL ↓ | Brier ↓ | ECE ↓ | |:---|---:|---:|---:|---:|---:| | [ Kev-4B](https://huggingface.co/jaredpalmer/kev-4b) | 74.19% | 75.32% | 0.858 | 0.380 | 0.125 | | [ Kev-27B](https://huggingface.co/jaredpalmer/kev-27b) | 82.31% | 85.28% | 0.533 | 0.265 | 0.050 | | [ Jev 1.13.0](https://docs.typesafe.ai/models) | 85.37% | 86.58% | 0.644 | 0.212 | 0.033 | | [ Laya](https://huggingface.co/convaiinnovations/laya) | 52.29% | 58.01% | 1.264 | 0.615 | 0.127 | | **JevAny releases** | | | | | | | [ JevAny-Gemma-4B](https://huggingface.co/SimpleJev/JevAny-Gemma-4B-LoRA) | 70.84% | 77.49% | 0.706 | 0.369 | 0.056 | | [ JevAny-Qwen3.5-4B](https://huggingface.co/SimpleJev/JevAny-Qwen3.5-4B-LoRA) | 78.68% | 80.09% | 0.587 | 0.297 | 0.035 | | [ JevAny-Qwen3.5-4B-Direct-Token](https://huggingface.co/SimpleJev/JevAny-Qwen3.5-4B-Direct-Token-LoRA) | 78.20% | 80.95% | 0.564 | 0.291 | **0.029** | | [ JevAny-Muse-Glimmer-30B](https://huggingface.co/SimpleJev/JevAny-Muse-Glimmer-30B-LoRA) | 83.46% | 87.45% | 0.464 | 0.229 | 0.032 | | [ **JevAny-Qwen3.8-27B**](https://huggingface.co/SimpleJev/JevAny-Qwen3.8-27B-LoRA) | **86.04%** | **90.04%** | **0.388** | **0.195** | **0.026** | NLL, Brier and ECE are measured on Transfer.

Full results and protocols · Machine-readable results · Method and ablation report

The public external comparison uses the same ten-model cohort on the complete 2,000-decision Typed Decisions test split and the same 724 JevJudge text requests. JevJudge uses native inputs without truncation under a shared 65,536-token context ceiling; every model answers 724/724. Typed Decisions measures agreement with teacher-derived soft gold, not objective correctness.

Two-panel comparison of the same JevAny and locally rerun open-model cohort on Typed Decisions and the JevJudge common text set

Model Typed Decisions ↑ JevJudge text ↑ Source / status
JevAny-Qwen3.8-27B 72.80% 66.44% Ours; full-context rerun
JevAny-Muse-Glimmer-30B 69.95% 62.57% Ours; full-context rerun
JevAny-Qwen3.5-4B-Direct-Token 67.20% 58.56% Ours; full-context rerun
JevAny-Qwen3.5-4B 63.50% 57.87% Ours; full-context rerun
JevAny-Gemma-4B 66.25% 51.10% Ours; full-context rerun
OpenDecider-small 66.65% 54.97% Open; full-context rerun
Bongard-mini 59.65% 51.24% Open; full-context rerun
Jeff-Gemma4-E2B 57.95% 46.69% Open; full-context rerun
Jeff-Qwen3.5-2B 55.45% 42.82% Open; full-context rerun
Jeff-Qwen3.5-0.8B 49.15% 43.51% Open; full-context rerun

Qwen3.8-27B leads the strongest locally rerun external model by 6.15 points on Typed Decisions and 11.46 points on JevJudge text. Direct-Token 4B leads by 0.55 and 3.59 points; Pointer 4B trails by 3.15 points on Typed Decisions but leads by 2.90 on JevJudge text, while Gemma-4B does not beat the strongest external row. Published-only Decider 1 (76.8%) and Liquid d1 (74.2%) remain ahead of our 27B on Typed Decisions and have no matching JevJudge result.

JevJudge text covers four roles and is a diagnostic, not the official five-role full-multimodal headline. Benchmark-trained specialist checkpoints are excluded from this zero-shot chart. Laya uses silent max_len=512 truncation and is not full-context comparable; Laya, Kev and published-only rows remain clearly labeled in the supplemental external tables.

Full external tables and reproducibility notes · Machine-readable chart results

⏱️ 3.1 Inference efficiency

On one H200, CUDA Graphs cut Qwen3.8-27B median latency from 113.54 to 30.53 ms (3.72×) on 231 JevBench questions; fused SDPA plus CUDA Graphs cut Muse-Glimmer-30B from 100.71 to 43.25 ms (2.33×) on a balanced 44-request Transfer panel. Accuracy stayed at 207/231 and 38/44, with no argmax changes. The table uses the H200 headline measurements for 27B and 30B, while retaining the original apples-to-apples A100-40GB comparison for the 4B models. Latency is comparable within each row; the fixed evaluation panels are listed explicitly. In the plot, diamonds show the 27B/30B H200 arrows and circles show the A100 cohort; the H200 points are not mixed into the A100 frontier.

Accuracy vs median latency before and after acceleration for JevAny and other decision models

Model Hardware Before After Speed-up Accuracy check Fixed panel
JevAny-Qwen3.5-4B A100 104.6 ms 25.3 ms 4.1× 78.68% → 78.87% Transfer, 1,046
JevAny-Qwen3.5-4B-Direct-Token A100 106.4 ms 25.9 ms 4.1× 78.11% → 78.39% Transfer, 1,046
JevAny-Gemma-4B A100 106.3 ms 31.9 ms 3.3× 70.84% → 70.84% Transfer, 1,046
JevAny-Muse-Glimmer-30B H200 100.71 ms 43.25 ms 2.33× 38/44 → 38/44 Transfer sample, 44
JevAny-Qwen3.8-27B H200 113.54 ms 30.53 ms 3.72× 207/231 → 207/231 JevBench public, 231

Median model-call latency, serial batch size 1. See the full report for the apples-to-apples A100 comparison and panel limitations.

Animation: Default, kernels with fused SDPA, and CUDA graphs race on one A100 clock slowed twenty times for JevAny-4B, 4B-DT and Gemma-4B; CUDA graphs finish at 25–32 ms, 3.3–4.1× sooner than Default

4B releases on one A100-40GB, Transfer-v9. Each bar fills at 1/20 of real time and stops at that stage's median latency.

Full tables, setup and other models · How to enable · H200 results · A100 results

🕹️ 4. Examples & Test Environments

The playground includes the three environments below. These GIFs preserve historical model actions and option probabilities; run the current JevAny-Qwen3.8-27B checkpoint with the commands in the playground guide.

🤖 4.1 Robot peg insertion

Use a Franka gripper to grasp, align and insert a peg, checked by PyBullet contact physics.

🔫 4.2 Doom corridor · 3D

Clear the final room by defeating the enemies on the left and right, then move forward. The environment uses ViZDoom and the included Freedoom assets.

⛏️ 4.3 Crafter survival · 2D

Gather wood, craft tools and mine stone while managing health and supplies.

🎮 4.4 Try the playground

With a model running from Run locally, open the playground:

jevany demo --base-url http://127.0.0.1:8008 --text-only

Open http://127.0.0.1:8090, choose Test and connect, and try your own decision. To let the model control a game, install the optional engines and restart the playground:

python -m pip install -e '.[demo]'
jevany demo --base-url http://127.0.0.1:8008 --text-only

Choose Run model, then One decision or Run automatically. Play yourself lets you control the game. Live control sends text state to the model; robot control uses the .[robotics] extra.

For the bundled recordings, run jevany demo and choose Replay. Playback works on CPU without model weights. See the playground guide for platform requirements and environment APIs, or integrations to combine JevAny decisions with an LLM planner.

🧩 5. Supported Model Families

Model IDs, supported inputs and setup requirements.

26 supported models across Qwen, Gemma, Muse, Mistral, GLM, Nemotron and Llama

📚 6. Documentation and Contributing

Training · Deployment · API · Data · Evaluation · Agent harness protocol · Contributing

To contribute a model adapter, evaluation or application example, start with the contribution guide. The technical report describes model design, multimodal support, the agent-harness study and appendix, negative results, and open questions.

Code and starter data are Apache-2.0. Some components are adapted from Kev; see NOTICE and ACKNOWLEDGEMENTS.md. Base models and upstream datasets retain their own terms.