Frontier AI research lab. We build AI scientists & the platforms they run on — compressing the full research lifecycle from idea to impact.

Robot-arm simulation + computer use: JEV-9B now sees. Now multimodal: image-based System 1 decisions join its calibrated text decisions and System 2 reasoning. New community-requested visual demos: shopping, settings, mail, and a MuJoCo robot-arm simulation. AutoTrust leads open decision models. Model ↓ huggingface.co/autotrust/JEV… For AutoTrust AI’s full research-agent workspace, download ScienceGuru for Windows or macOS: scienceguru.ai/
Thanks @Gorden_Sun, @potechi_takusan, @Read_Acted, and @a_captaincook — the questions have been great. A few clear answers: • Computer use: yes. The path from screen state → calibrated next action → click / type / scroll loop is a core direction for JEV-27B-VL. • 3D + robotics: also yes. The current release takes 2D image state; 3D scene understanding and embodied / robotics workflows are directions we are actively building toward, not capabilities we claim are fully solved today. • Local deployment: agreed — an 80GB GPU is too much for many builders. The current BF16 release targets that class of hardware, but quantized and GGUF variants are coming soon. • For agent engineers: JEV-27B-VL is open-weight under Apache-2.0, with vLLM and SGLang deployment paths. It is designed for calibrated visual decisions, not just image description. Mario and Tetris make the loop easy to see: see → decide → act. The destination is broader: computer use, visual agents, and eventually embodied systems. JEV-27B-VL: huggingface.co/autotrust/JEV… For AutoTrust AI’s full research-agent workspace, download ScienceGuru for Windows or macOS: scienceguru.ai/
5
249
Most decisions don't need a reasoning model. The hard ones do. GEV-26B-Decide knows the difference. It answers in ~45 ms, and stops to think only when it's unsure. What a moment of thought buys: 💣 Minesweeper: 21% → 86% 🔴 Connect Four: 54% → 99.5% 🟩 Wordle: 52% → 100% 🔢 Sudoku: 76% → 99.5% JEV Decision Index: 53.3 for our JEV-27B-VL → 62.48 for GEV-26B-Decide, even higher than the current leader, Cloudflare's Clef (61.2). All self-reported. Open weights 👇 huggingface.co/autotrust/GEV…
2
8
16
1,062
AutoTrust AI Lab, in motion. JEV-9B, JEV-27B, JEV-27B-VL & JEV-Gemma4-26B-A4B pair fast probabilistic decisions with reasoning. GEV-26B-Decide adds adaptive thinking. SLIM-Q GLM-5.3/Flash GGUF builds target DGX Spark. Watch the models play ↓ huggingface.co/autotrust
September at AutoTrust AI Lab: 6 open-model releases—JEV-9B, JEV-27B, JEV-27B-VL, JEV-Gemma4-26B-A4B + two GLM-5.3 GGUF builds for DGX Spark. Watch JEV-27B-VL tackle Mario, Tetris & Rubik’s Cube. Bonus: an October GEV-26B-Decide preview.@AutoTrustAI huggingface.co/autotrust
2
8
590
Thanks @Gorden_Sun, @potechi_takusan, @Read_Acted, and @a_captaincook — the questions have been great. A few clear answers: • Computer use: yes. The path from screen state → calibrated next action → click / type / scroll loop is a core direction for JEV-27B-VL. • 3D + robotics: also yes. The current release takes 2D image state; 3D scene understanding and embodied / robotics workflows are directions we are actively building toward, not capabilities we claim are fully solved today. • Local deployment: agreed — an 80GB GPU is too much for many builders. The current BF16 release targets that class of hardware, but quantized and GGUF variants are coming soon. • For agent engineers: JEV-27B-VL is open-weight under Apache-2.0, with vLLM and SGLang deployment paths. It is designed for calibrated visual decisions, not just image description. Mario and Tetris make the loop easy to see: see → decide → act. The destination is broader: computer use, visual agents, and eventually embodied systems. JEV-27B-VL: huggingface.co/autotrust/JEV… For AutoTrust AI’s full research-agent workspace, download ScienceGuru for Windows or macOS: scienceguru.ai/
Introducing JEV-27B-VL. We believe it is the world’s first open-weight, near-SOTA multimodal decision model — built with AutoTrust AI’s Blocks of Experts (BoE) recipe. It does not just describe what it sees. It turns visual state into calibrated action probabilities, then selects what to do next. In this demo, it sees Mario’s position, movement, obstacles, and timing — then chooses whether to run, jump, or do both. See → decide → act.
1
20
1,975
Here is the mechanism. 1. System 1 returns a calibrated distribution over the options in one pass. 2. If its top option is below the confidence threshold — 0.8 by default — System 2 reasons over the same state, question, and choices. 3. The final answer combines the fast System 1 distribution with the reasoning-informed distribution. The goal is not “always think.” It is to spend compute where it can actually change the decision. The trade-off is measurable. Across six held-out reasoning sets, adaptive mode reached 83.4% accuracy versus 73.3% for System 1 alone. It triggered thinking on 47.8% of questions, while capturing 96% of the gain from always thinking. For routine decisions, System 1 answers in about 45 ms on one B200. On our Decision Index 0.2.1 recomputation, GEV-26B-Decide reaches 62.48 balanced skill. This is a self-reported run, not a leaderboard entry.
1
20
GEV-26B-Decide supports calibrated: • yes/no decisions • 0–5 scoring • choices over 2–256 options • text or image state • prompts up to 256K tokens The model card includes the videos, reproducible reports, a vLLM server, and the /v1/decide API. Model + code: huggingface.co/autotrust/GEV… For an AI-native research workspace from AutoTrust AI, download ScienceGuru for Windows or macOS: scienceguru.ai/
1
1
31
AutoTrust retweeted
JEV-27B-VL:开源多模态版本Jev模型 像人类一样有两种思考模式: 第一种是“直觉反应”(快思考)。面对选择题、判断对错或者打分这类任务,它看一眼文字或图片,不用逐字生成长篇大论,一次计算就能直接给出每个选项的概率。比如判断一张照片是不是食物,或者在几十个分类里挑出最合适的一项,反应速度只需要几毫秒到几百毫秒。 第二种是“深度分析”(慢思考)。当遇到复杂的图表或需要深入推导的逻辑题时,它会切换到传统的分析模式,一步一步梳理细节并输出完整的文本解释。 支持图片输入,能用于视频推荐、玩游戏等场景。 模型:huggingface.co/autotrust/JEV…
2
6
54
6,820
Thanks, @mr_r0b0t — really glad you enjoyed it! 🤩 The Mario demo is exactly the kind of loop we wanted to make tangible: visual state in, calibrated action probabilities out, then the agent acts. JEV-27B-VL is open-weight under Apache-2.0, so we’re excited to see what people build with it.
Open weight decision model with vision!? Yes please 🤩 Pretty wild watching it play Super Mario 👀 huggingface.co/autotrust/JEV…
1
372
A week of JEV at AutoTrust AI: from fast text decisions to open multimodal action loops. JEV is our System 1 family for agents. Instead of generating a paragraph before every small action, JEV returns calibrated probabilities for yes/no, 2–16-way choice, and 0–5 scoring — so an agent can act immediately, or escalate hard cases to System 2 reasoning. The foundation is our Blocks of Experts (BoE) recipe: keep a capable base model intact, add a small decision expert, and route each request to the right path. Fast decisions and deeper reasoning, served from one engine. What we have released: • JEV-9B — our first integrated open System 1 + System 2 model. It reaches ≈0.019 KL to TypeSafe Jev 1.13’s decision distributions, while making a decision in about 90 ms median on one B200. Its System 1 block trains only 40.2M parameters — 0.5% of the backbone. • JEV-27B — our stronger 27B model. It reaches ≈0.017 KL, achieves 96% of Jev’s accuracy on an independent 16-option benchmark, and scores 84.07% across six public decision benchmarks in our updated evaluation, ahead of Jev 1.13’s 83.85% in 4 of 6 groups. Its System 2 path remains unchanged at 78.0% HumanEval, while System 1 runs at 137 ms median per decision on one B200. • JEV-Gemma4-26B-A4B — a BoE System 1 + System 2 bundle on Gemma 4 MoE. It scored 58.05 on Decision Index 0.2.1, completing all 150,317 scored requests with 0 errors and 0 unsupported. • JEV-27B-VL — what we believe is the world’s first open-weight, near-SOTA multimodal decision model. It extends calibrated decisions to images and visual state: Mario, Tetris, and Rubik’s Cube are not chat demos; they are see → decide → act loops. Across the JEV family, the goal is the same: give agents a fast, calibrated System 1 for high-frequency judgments — without giving up the System 2 reasoning needed for hard work. Explore the JEV family, models, demos, and evaluation results: huggingface.co/autotrust Then use these capabilities inside a full research workspace. Download ScienceGuru for Windows or macOS: scienceguru.ai/
Introducing JEV-27B-VL. We believe it is the world’s first open-weight, near-SOTA multimodal decision model — built with AutoTrust AI’s Blocks of Experts (BoE) recipe. It does not just describe what it sees. It turns visual state into calibrated action probabilities, then selects what to do next. In this demo, it sees Mario’s position, movement, obstacles, and timing — then chooses whether to run, jump, or do both. See → decide → act.
2
1
9
1,395
Introducing JEV-27B-VL. We believe it is the world’s first open-weight, near-SOTA multimodal decision model — built with AutoTrust AI’s Blocks of Experts (BoE) recipe. It does not just describe what it sees. It turns visual state into calibrated action probabilities, then selects what to do next. In this demo, it sees Mario’s position, movement, obstacles, and timing — then chooses whether to run, jump, or do both. See → decide → act.
70
158
1,534
151,718
JEV-27B-VL is open-weight under Apache-2.0. The updated model card includes: • Weights and decision adapter • BoE architecture and System 1 / System 2 setup • Calibration details and decision prompts • vLLM and SGLang serving instructions • Image recommendation evaluation • Local deployment examples Download the model and run it locally on a GPU with 80 GB+ memory: huggingface.co/autotrust/JEV…
4
10
67
3,653
JEV-27B-VL is built for the fast visual decisions inside an agent loop: what to inspect, what to click, what to retry, what action to take next. ScienceGuru is where those capabilities connect to real work — research projects, files, code, terminal, experiments, long-term memory, and multi-step workflows in one desktop workspace. Download ScienceGuru for Windows or macOS: scienceguru.ai/
2
15
2,837
Decision models are becoming core infrastructure for agents. JEV-27B is our open-weight System 1: calibrated yes/no, 2–16-way choice, and 0–5 scoring in one forward pass — 137 ms median per decision on a single B200, measured server-side without network latency. Today, we add vision with JEV-27B-VL — an open-weight multimodal decision model built with AutoTrust AI’s Blocks of Experts (BoE) recipe. It makes calibrated yes/no, 2–16-way choice, and 0–5 scoring decisions directly over text and images, while keeping System 2 reasoning in the same engine. On MicroLens-100k, its zero-shot, cover-only recommendation reaches 0.727 AUC: on par with collaborative filtering trained on 59,045 users’ watch histories, with a higher HR@5. Updated model card: huggingface.co/autotrust/JEV…
OpenAI’s just killed to TypeSafe’s Jev System 1 model OpenAI just shipped a 150ms decision engine. built for agents that need to pick the next move, not write a paragraph about it. Give Luna text or an image + a fixed set of answers. Get the choice back. OpenAI’s new Decisions API (Luna) takes context + a closed list of options and returns the decision in 150ms. Instead of wasting tokens on an explanation you’ll throw away, you get the decision directly.
1
273
AI agents live or die by the small decisions: select, route, verify, act. We’ve released the complete Decision Index 0.2.1 outputs for JEV-27B on @huggingface. • 53.30 Decision Index score • 64.32 raw score • 150,317 / 150,317 scoreable requests completed • 0 errors, 0 unsupported requests This is more than a headline number. The open result set includes compact answers, probabilities, timings, statuses, environment records, engine code, hashes, and training-overlap checks—so the run can be inspected and re-scored with the Decision Index kit. JEV-9B results are included too: 43.14 → 53.30 on the same Decision Index edition. Explore the open results: huggingface.co/datasets/auto…
Open source matters more when people can actually try it. Huge thanks to @huggingface and @multimodalart from the Hugging Face open-source team for building a live interactive JEV-9B demo on Spaces, powered by ZeroGPU. JEV-9B is our open System 1 model for calibrated, structured decisions. Try the demo: huggingface.co/spaces/autotr… huggingface.co/spaces/autotr… Model card: huggingface.co/autotrust/JEV… huggingface.co/autotrust/JEV…
2
3
243
What happens when one AI research system is tested across four public RSI testbeds in a single month? Thank you @MorningstarInc for featuring our release on AutoTrust AI and ScienceGuru. In September, ScienceGuru: • reached #1 on the official Autoresearch@Home leaderboard • posted the top maintainer-validated NanoPath v2 entry when submitted • reported the fastest times on two GPT-2 training tests (self-reported; maintainer review pending) The common thread: one system reading prior work, proposing ideas, writing code, running experiments, checking results, and publishing the evidence. Read the full story: morningstar.com/news/pr-news…
1
2
3
134
Try ScienceGuru for Windows or macOS: scienceguru.ai/
89