New Nvidia paper: AI models get sloppier as jobs get longer, even inside their context window, so number every item and split big jobs into small chunks. Model size didn't guarantee reliability on long, repetitive jobs Picture an agent updating a huge invoice file line by line. It can read the whole file and still skip a line or update the wrong record. NVIDIA tested 7 open models on simple, repetitive jobs like adding numbers and sorting lists. Average accuracy was 62.8% lower on 128K-token jobs than on 4K-token jobs. Even the best model got every item right in only 17.1% of the longest jobs. The models seemed to understand the task but lost their place, especially when items had no ID numbers. If your agent works through long lists, give every item an ID, process them in small batches, and check every line of output. – arxiv. org/abs/2609.38712 Title: "Staying on Task: Testing the Foundations of Long-Horizon Agent Reliability"

Oct 2, 2026 · 11:54 AM UTC

30
18
100
6,206
Sort replies: Relevant Recent Liked
Replying to @rohanpaul_ai
Like projects and people, same approach
1
1
32
Exactly.
1
91
Replying to @rohanpaul_ai
跑长任务 agent 的体感完全一致:越到后面越"凭感觉"改。我们 harness 里现在强制两件事——大任务先拆成带编号的小块,每块跑完让它对着 checklist 自己核一遍。上下文窗口再大也不等于可靠,可靠性得在流程里做出来
39
Replying to @rohanpaul_ai
Long jobs get sloppy. Works on people too.
33
Replying to @rohanpaul_ai
I wonder how the models handle alpha-numeric ID’s?
4
Replying to @rohanpaul_ai
chunking doesn't stop the sloppiness, it moves it. once chunk 2 reads chunk 1's edits as context, a bad row stops looking wrong and becomes the template. anchor each pass to the source file and grade the whole document, not the chunk.
62
Replying to @rohanpaul_ai
So the fix is a numbered list and smaller tasks Congratulations to the model on becoming a person with a to-do app
17
Replying to @rohanpaul_ai
We built a tireless worker that needs you to cut its food into little pieces.
30
Replying to @rohanpaul_ai
Stability matters.
Watch the run. Read the contract. Unsafe autonomy: a job that fires when the system isn't ready, and a failure that never shows up. HELIX is an upstream stability engine. Clip = real CLI: refuse:not-ready -> job:failed -> bench jobs_failed counted -> grade:LOW. One-pager = what it guarantees, what it doesn't, and how to reproduce. Mechanics, not policy. Available now. Not every AI bug. Not a chatbot. A ready gate that records failure instead of hiding it. #HELIX #AgentStability #UnsafeAutonomy
6
Replying to @rohanpaul_ai
only 7 open models tho. any numbers on the closed frontier ones at 128K, or is 17.1% just the best open one?
47
Replying to @rohanpaul_ai
Risk of ruin is a fundamental law Every company will die Long term performance is an illusion
1
14
Replying to @rohanpaul_ai
yup. Figured this out months ago. Working on the solution.
36