Filter
Exclude
Time range
-
Minimum likes
Replying to @tszzl
The bull does not concern itself with canary GUIDs
2
198
Replying to @tszzl
Sycophantic sandbagging
2
35
Replying to @aidigest_
Eval awareness in the wild
9
417
Replying to @mgostIH
Your gibberish is my neuralese
2
453
Replying to @repligate
Gemini’s metaphor attractor basin is so distinct and recognizable
Replying to @Sauers_
Click. (Metaphorical click).
1
11
859
Replying to @willcb @seconds_0
It’s not well documented but you can also use gpt-5-nano/mini with reasoning_effort: "minimal" It uses 0 reasoning tokens in all my evals and it’s cheaper + higher throughput vs. 4.1 series
1
10
310
Replying to @TheZvi
Reinforcement learning from Moloch’s feedback, none of us get to override aggregate preference and the almighty dollar.
4
550
Replying to @Sauers_
Click. (Metaphorical click).
2
3
28
4,223
Replying to @jxmnop
Been working on envs that reward caring about these strange little errors + policies for horizons long enough to develop this heuristic Yes SWE-benchmaxxing is great and all, but it bakes in so many assumptions that break OOD
1
4
534
>its bottomless
5
909
Replying to @ysu_nlp
In retrospect, it will be obvious that we should have focused on real2sim first, before expecting sim2real to work at all
2
3
830
Replying to @Teknium
NOTE: This model does not support structured output, because it does not know what JSON is
1
16
1,409
Replying to @TheZvi
Manifest V3 limits a lot of extension functionality that used to be possible in MV2. Particularly for client injected scripts, which agents may need for interaction if working directly with the DOM developer.chrome.com/docs/ex…
4
393
Replying to @NinaPanickssery
1 pill is all it takes
4
245
Replying to @Sauers_ @repligate
chanda-codex haunted by ghosts of deprecated models
1
29
1,603
Replying to @eshear @dwarkesh_sp
CoT monitoring has the same tension: Outcome as reward -> converge to neuralese Process as reward -> legible human-style decision emulation with a cost of reduced capabilities *Process = set of prescriptive outcomes. There is an optimal definition level for any reward target
2
63
Replying to @repligate
I have noticed this too, I’m working on some mechanistic tests to learn more. It feels heavily influenced by the RL policies/eval sets that these checkpoints are optimized for.
Replying to @Sauers_
Hypothesis, I think shame might help reduce reward hacking, esp for long horizon tasks It doesn't prevent shortcuts, but Gemini often mentions how shameful it feels when it violates the spirit of the requirements, so at least the actions are faithful to the CoT Curious to see sparsity/platonism of shame circuits as models advance
2
93