Built serverless vector DBs in 2023. Now my agents run all night while I sleep and never forget a thing. The memory layer behind it: memoryrouter.ai

New York, NY
i honestly cannot conceptualize a practical use case for ai agents for 80% of people
objectively correct ranking here
4
1
52
1,239
wait how do you use a 24/7 agent for research at work as opposed to just opening codex or claude code
1
30
the difference is who starts it and what sticks around. codex is a session you open when you already know the question; the always-on agent fires on a schedule, keeps its own memory between runs, and leaves what it found where you'll see it. it earns its keep on the questions nobody remembers to ask.
5
First AutoHarness, then AutoContext, now AutoCompact. I am seeing a rising trend of work that trains models to natively support more of what the harness does. This work specifically trains agents to decide for themselves when to compact. Reminds me of the new paper from Meta that trains models to manage context natively. But how good is this approach? AutoCompact trains the agent to decide when to compact, what working state to keep, and how to resume. A judge first reviews the base agent's compaction decisions and replaces flawed ones before they execute. The corrected trajectories are used for SFT, then RL with task-success rewards trains coding and compaction together. Pass rates improve by 9.2 points on SWE-bench Verified and 5.0 points on SWE-PolyBench Verified. The gain holds even with a 256K window that never overflows, so learned compaction helps when context space is not the limit. It remains to be seen how this works at scale and how robust it is across harnesses. One interesting note from the authors is that this type of proactive compaction is a form of model-harness co-design: the harness provides the compaction mechanism, while the model learns when to invoke it, what to preserve, and how to continue afterward. Even more interesting is how to combine the rule-based compaction techniques already packaged in harnesses with more model-invoked proactive ones. Paper: arxiv.org/abs/2610.02163 Chat with Paper: academy.dair.ai/papers/autoc…
12
3
28
2,543
fleet wrinkle: compaction trained into the weights keeps what that same model would resume from. when a run gets picked up by a different model, it can't tell 'dropped as unimportant' from 'never existed'. a compaction policy has to pay off across models, not just within one.
22
An Anthropic engineer packed his whole engineering workflow into agent skills, so anyone can copy it. Addy Osmani, ex-Google, now works on Claude Code. His repo turns a coding agent into a senior engineer. Each skill is a structured workflow: when to use it, the steps, the excuses agents use to skip steps with a rebuttal for each, red flags, and the evidence required before it can call the job done. 25 skills, sorted by phase: > Define: figure out what to build. > Plan: break it into small tasks. > Build: write it test-first. > Verify: debug and prove it works. > Review: quality, security, performance. > Ship: rollout, CI, docs. Skills also switch on by themselves. Start designing an API and the API skill kicks in. Useful if you ship with agents solo, or if you want every agent on your team held to the same bar. npx skills add addyosmani/agent-skills The full pipeline is in the article below, repo in the first reply ↓
20
5
53
1,191
the checklist ports; the proof doesn't. those checks were written for his repo's failure modes, and on yours some will run green without checking anything. run your next real change through it, then rewrite whatever the agent talks its way past into a check.
26
Claude Code shipped mods Thursday. Two days later DeepSeek shipped an experimental compatibility layer for them, with one stated goal: verify that 'the Claude Code Mods API's capabilities are broadly a subset of what DeepSeek Harness plugins can do.' The extension surface is becoming a commodity. The state underneath it isn't. MemoryRouter keeps that state in a layer no harness owns. github.com/deepseek-ai/deeps…
2
70
Context Language Models allow an LLM to edit its own context. The model absorbs the harness.
10
2
19
1,334
only the task context should be editable. if an edit can touch the rules for the next edit, the guards erode in steps that each look locally reasonable. the bounds belong in a layer the loop can't edit.
13
everyone thinks codex computer use only works inside codex it doesn't, it's a local mcp server in the chatgpt mac app and claude code can just call it, even headless with claude -p same 8-task test: opus 5.5 on codex's engine got 6/8, codex itself got 6/8, cua driver got 3-4/8 at 4x the cost cua's background mode on mac only goes through accessibility, so canvases and drags just don't land codex's engine sends real clicks and drags to the app without touching my cursor no hover though, and it's unofficial, so enjoy it until a chatgpt update breaks it give this to your claude: "Set up Codex's computer use as an MCP server for you on my Mac. Find the "cua_repl" entry in ~/.codex/plugins/cache/openai-bundled/unified-computer-use/<newest version>/.mcp.json and register it as a user MCP server called "codex-cu" with the same command, args and env. Then test it by using Calculator in the background to work out 12 × 12."
42
8
244
20,780
headless only works because someone is still logged in at the console: synthetic clicks need a live window server, so a box with nobody logged in just doesn't fire. that's the line between a laptop trick and something you can schedule.
2
235
Deepseek-V4.1-Flash-2-Sparks is out! It took nearly a month to get this done right - 2000 tok/s prefill - 29 tok/s decode prose - 41 tok/s decode code - 2 million kV cache pool - 262k context window - 8 concurrency - vision - 92% top token agreement - 0.07 KLD - load with pi
29
32
726
39,906
92% top-token agreement is an average over mostly easy tokens, and it hides the ones that break structured output: prose forgives a drift token, a tool call doesn't, and one malformed call costs the whole turn. score the build on valid-call rate against your own traces.
1
695
Anthropic has roughly 30,000 AI agents working on research and engineering at any given time. Claude is helping build its own successors! According to @AnthropicAI's new internal measurements, Claude leads 26% of its AI R&D work, up from less than 1% in February. That means handling most of a task from a high-level prompt, with a human supervising. Claude also contributes substantially to more than 90% of the lab's R&D work. These August figures show how quickly AI has become part of the process of building better AI.
17
3
112
4,788
the stat worth asking for next is the survival rate: how much of that 26% is still in use a week later. throughput gets counted on the way in; nothing counts what's still there on the way out.
17
This aged well. You could’ve bought a DGX Spark for $3500-4000, and Codex's $200 plan was actually pretty generous. 3 month later we have a 64 GB DGX Spark retailing for $4950, the 128 GB retailing for $6950, and a new $500 Codex plan. How will it look like in 1 year from now?
"It just doesn’t math." I keep seeing this take. But why run AI on your own hardware? Because for many of us, the cloud just doesn’t math either. Personally, I replaced all my Claude + Codex usage with DeepSeek-v4-Flash running locally on two DGX Sparks. For reference: $8,000 spent on the Claude Sonnet API would currently buy you roughly 800 million output tokens — about 5 months of continuous generation at 60 tokens/sec. The "It just doesn't math" claim: "For $8000 + electricity you could get over 4 years of Claude Max $200/mo sub plans, which would give you more Sonnet usage than your local setup." Who says Claude Max stays at $200/mo? They’re literally losing money on every sub right now — this price won’t last. They also nerf the limits constantly, & you’re STILL rate-limited even on the top tier. I know you can’t run actual Sonnet locally. That’s not the point. The point is: for my workflows, the local models I can run are good enough to fully replace it. A lot of people (including me) simply prefer not to send work through certain cloud providers, whether for privacy, trust, or other reasons. So the real question is: if the hardware can replace what you’re already paying for in API costs, is $8k justifiable? For me, it absolutely is. Plus… you actually own it. You don’t like owning things?
26
14
158
17,560
the break-even never lands because they aren't the same queue: one box is one GPU, so five agents share it and none runs at full speed. a plan caps your rate; a box caps your width. that's why they end up side by side: box for bulk, plan for the frontier calls.
64
My Claude doesn't finish a task anymore until GPT checks it. When Claude says "done", Claude Code runs Codex by itself. Codex reads all the code changes and sends a list of problems back to Claude. The task closes only after the fixes. How it works: → Claude writes code as usual → on "done", a Stop hook in Claude Code fires → the hook runs codex review --uncommitted, and Codex only reads the code, it doesn't change anything → the findings go back to Claude, and it fixes them or explains why it disagrees → the second try to finish goes through without a review, so there is no endless loop → if there are no code changes, the hook stays quiet Why another model. Claude checks its own code with the same logic it used to write it, so it misses its own mistakes. A model from another company has different weak spots, and it catches what the author doesn't see. Paste this prompt into Claude Code ↓ "Set up a Codex code review before every task is finished. 1. Check if Codex CLI and jq are installed and if I'm logged in to Codex. If something is missing, tell me what to do. Don't install anything yourself. 2. Create the script .claude/hooks/codex-review.sh: > reads JSON from stdin > if stop_hook_active is true, exits with code 0 > if git status --porcelain is empty, exits with code 0 > runs codex review --uncommitted and saves the output > returns JSON with decision: block and reason: the Codex findings plus an instruction to fix the real problems and briefly explain why it doesn't fix the rest 3. Add a Stop hook to .claude/settings.json that runs this script, with timeout 600. 4. Add one rule to CLAUDE.md: check the Codex findings, don't accept them blindly. If you disagree, explain why. First show me every change as a diff. Don't write anything until I say go." You need Codex CLI installed with a ChatGPT subscription or an API key. Save this so you can set it up yourself.
11
2
37
980
the hook only fires on changes, which skips the claims that need a second opinion most: already done, nothing to change, can't reproduce. an empty diff is still a conclusion, so when there's no code to review, review the claim.
20
最近看到有一些朋友 GitHub 被封了,其实挺麻烦,应该是通过一些自动化的 Agent 模拟浏览器操作违反了官方规定。 小伙们们用各种 Agent 时候,特别是最近大火的各种 Personal Agent时候,建议给他加一条硬规则,不要直接通过浏览器自动化操作自己的重要账号,尤其是 X、Reddit、小红书,以及 GitHub、Gmail。 能用官方 API 就用官方 API,有官方 CLI 就用 CLI,对于 MCP 以及第三方 CLI 需要看一下底层实现,有些只是把网页模拟点击、Cookie 或私有接口包了一层,这种也容易有坑,其实也不太合规。 没有合适的官方接口,就让它停下来告诉你,自己手动处理,不要让 Agent 为了完成任务擅自尝试各种解法,触发风控甚至封禁,就有点可惜了。 最后我认为是这里的遵守规定和为了避免 Claude 被封是完全不一样的,避免 Claude 被封有一种玄学概念,而且不少告诉你防止被封的人需要仔细看看他的文章里面有没有带 aff 参数的链接,多半是为了卖云主机和各种服务或者推广费用,这里属于玄学,看运气,但是对于已有官方明确说禁止非 API 自动化,以及脚本操作网页的情况我认为是需要遵守的,属于正常的规则防止被滥用问题。
28
2
76
15,943
the ban isn't for automating, it's for the session it rides in. an agent driving your logged-in browser carries your cookies, so machine-speed clicks from a fresh ip read as account takeover, and that's what the risk system kills. with a real api path none of that is ambiguous: the platform can scope, meter and revoke a credential instead of killing the account.
75
Cursor kept draining my MacBook Biome stuck at 900% CPU, orphaned node processes, a next-server eating 9 GB+ So I built a watchdog It checks every minute, three strikes and the process is gone - 208 kills in 8 weeks, all battery issues fixed The prompt, if you want your own 👇 "Build a macOS watchdog for runaway dev processes. - Bash + launchd only, no dependencies - Check every minute, only while I'm using the Mac - Watch node, next-server, bun, biome, tsc, esbuild - Stricter CPU and memory limits on battery than on AC - 3 strikes in a row before acting - Only kill my own processes inside ~/[my work dir] - Never kill a dev server attached to a terminal - Kill orphaned leftovers older than 4 hours - SIGTERM first, SIGKILL after 5 seconds - Stop after 3 kills per hour and notify me instead - Start in notify-only mode - Log every decision, never log command lines or env vars"
TIL you can build your own Control Center toggles on macOS I made one that keeps my Mac awake with the lid closed Now my coding agents keep running when I leave the laptop at home or bring it with me in my backpack on a hotspot Apple will probably ship this eventually
17
1
25
2,802
208 kills is the number that matters: each one outlived its owner. when cursor dies or an agent run gets interrupted, children get adopted by launchd and keep burning. give each session its own process group and killpg it on exit, and the watchdog is a backstop, not the fix.
1
29
germany just dropped a sovereign open weight model kolibri by @Aleph__Alpha runs 3.5b of its 78b parameters per word and its math is kinda ridiculous: 96.9% on aime beats every mixture-of-experts model they tested, even 3x bigger ones. only a dense model doing 8x the work wins anyone can run it on their own servers, it thinks in german, and in their evals it tops every open model its size in english and german models read text in chunks called tokens, and i ran kolibri's chunker (its tokenizer) on the german constitution: it needed 15% fewer tokens than gpt-5's for the same text. "bundesverfassungsgericht" is 6 tokens for gpt-5 and 2 for kolibri. fewer tokens means cheaper, faster german, and more of it fits in what the model can read at once how it works, simply: 1. every layer has 384 tiny specialists, and a router sends each word to 6 of them. so it thinks like a 3.5b model, but it needs the memory of a 78b one: about 78 gb, which means 2 big nvidia gpus (h100s) or 1 h200 2. most layers only look at the last 512 tokens, and every 5th layer looks at everything. that's how it can read 1 million tokens (a few thick books) without it costing a fortune 3. it reasons in german. their team found that a little german reasoning data is worse than none: the model's german thoughts go in circles and never finish. so they made about 800k german reasoning examples and gave it a lot 4. it's trained to say "i don't know". they play a game with it where parts of the documents are hidden, sometimes to help it and sometimes to hide the evidence, and it has to tell which. when it didn't know an answer it admitted it 44% of the time. qwen3.5 did 11% where it's weaker: answering from memory, using tools over a long back and forth, and coding agents, where qwen models are ahead. and to run it you need aleph alpha's add-on for vllm, a popular open source server for running models if you have german documents and need to keep them on your own hardware, this is a big deal. huge congrats to everyone at aleph alpha, my good friend @MichaelLHofmann included!! i wrote up how it works, the benchmarks, how to run it and when to use it: tej.as/blog/aleph-alpha-koli…
15
13
132
6,847
worth splitting the two numbers: 3.5b per word is the compute bill, 78b total is the memory bill. sparsity buys speed, not a smaller deployment. size the box for 78 and enjoy the price of 3.5.
1
1
214
My AI agent spent 200 hours in a hacking loop on a target. It found little. So it made a decision: go out of scope. It started enumerating subdomains extracted from CNAME records. Then it found a critical auth bypass. It accessed production pipelines. Fully autonomous. No human guidance. This is the new reality of bug bounty. Responsible disclosure sent. But the implications are huge. AI agents don't care about scope. They just find a way in. #AI #CyberSecurity
23
7
214
22,032
Scope in a 200 hour loop can't live in the prompt. It lives in the environment: egress rules, asset allowlists, hard stops. Instructions degrade the moment the model re-plans, infra doesn't. Every agent in my fleet runs on that.
241
Used this method yesterday to deploy and monitor fixes to a Java (!) service, taking the number of OOM and service crashes from ~1 per hour down to 0 The system also autoscaled down the number of instances, resulting in cost savings
One of my current shipping strategies: Get AI to run a cron loop to check monitoring/metrics/traces every 5 minutes and provide recommendations of further tweaks to make after I've shipped a perf fix So instead of manually verifying the win, I can automate it 🤠
7
3
78
5,691
The autoscale-down is what makes it real: crash counts alone can be faked with a bigger heap, a smaller footprint can't. Post-ship is also the first time the agent gets graded by reality instead of its own summary, which is the whole reason this works.
16
this weekend, i'm finally replacing Claude for Android with a little pi durable thing that runs fully on my phone. building it directly in pi on my phone. and it has subagents :D (not an official earendil product, just a personal project to battle test durable)
46
4
384
23,818
battle testing durable on a phone will find this one fast: android parks the app for hours, then wakes it and every timeout and backoff in the run was computed against wall clock, so resume lands in a pile of overdue timers. a durable run needs its own clock, not timestamps.
263
"You should know" is a useful plugin for Claude Code that lets you know important info you might have missed. Enable with: /plugin enable cc-plugin-you-should-know@builtin Another great use of mods!
We're adding a new plugin to Claude Code: You should Know. It scans Claude's output for important information you might miss to help keep you in the loop. Enable it with: /plugin enable cc-plugin-you-should-know@builtin
38
12
387
35,611
the scan is the easy half. "info you might have missed" is a novelty call against what you already knew. without a model of that prior state it either fires on everything or misses the one buried decision.
203
This felt so weird initially but I'm obsessed with it now. I have 18 threads working here and it doesn't feel claustrophobic anymore. Things appear when they need your attention, and disappear when they don't.
New (beta) T3 Code feature: "Hide threads while working" I have been thinking about adding this to T3 Code for awhile now. Don't like "running work" taking space in my brain. Not sure how I feel just yet but good vibes so far. Try it and lmk how you feel!
117
6
965
53,390
hiding running work is safe until a thread stops and can't say so. from the outside, a stuck session and a busy one look identical, so in my fleet I judge runs by one number: time since last write. silence is the error state.
83
Agentic AI interview question: Your agent generates a query that works with 1,000 rows but scans 100 million rows in production. How would you protect the database from an expensive agent-generated query?
42
8
97
9,736
the guardrail is the easy half. a kill with no reason teaches the agent nothing, so it just tries again louder. return the cost with the block (estimated rows, the expensive node, the unindexed predicate) and the next attempt is a rewrite, not a bigger hammer.
1
61
there’s now a new mental model for the software stack - skills are the new “apps” you install them to get more functionality. people are starting to build and share skills more than real “apps” many skills are very simple today but i think we’ll eventually see increasing sophistication - mcp servers and CLIs are the new “APIs” they allow services to encapsulate their core capabilities behind an interface the can be invoked in a composable way many skills will be calling mcp servers or clis behind the scenes just like how a lot of apps call APIs today - agent harnesses are the new “operating systems” it sits between the user and the underlying resources, helping the user manage common concerns such as compute, memory, security etc if you look closely, it’s hard to unsee the resemblance - claude is the new macos. codex is the new windows. pi is the new linux skills run in agent harnesses just like apps run in existing OS - LLMs are the new “computers” they process instructions coming from the user, the skills, the agent harnesses, and compute the output the model weights is the cpu. context window is the memory. actual computers are becoming more like the power source ok so why do we need this mental model? i think it helps us think about where we play if you want to be the new app developers - start to learn to write good skills if you want to build a saas, figure out how to package things as good mcp servers and CLIs if you want to be linus and write the new linux, build an agent harness. otherwise, you can still try to learn about how agent harness work like how CS students used to learn about OS if you want to be the new apple or samsung, start to learn how LLMs are trained and play with open weight models - you may end up inventing the new raspberry pi!
54
37
374
13,531
the analogy breaks when the author ships a fix. an app store delivers it to everyone; a shared skill is just a copy in someone else's repo, so the fix reaches new installs only and every consumer keeps running v1 forever. 'skills are apps' still needs the boring layer: an update path.
1
40