Context Manager. Writing about software and AI. Intuit. Ethos Life. A few other places. Egg avatar until I find a picture of me less than 700kb.

West Coast
How I Plan an Engineering Project End to End Caveat: This is how I operate at a big tech company. This would look different at a startup and vastly different if I was just building by myself. At a large company there is definitely a coordination tax you need to pay, and that seeps into how you build. 1. Build Context - Understand I usually spend the first couple of days just building context. I use an LLM to explore the codebases this might touch, read the PRDs, and understand how the project fits into the existing ecosystem. The goal is to build a strong mental model of what is already there and how well it is working. From there, I start figuring out which services we can extend, what might need to be new, and reaching out to the owners of the services that we might be extending. I will then write a couple of different design docs for the proposal. One high-level doc for the business that focuses on the why, and one low-level doc for the engineering team that focuses on the how. 2. High-Level Doc - Align For the high-level doc, the audience is generally very senior engineers/PMs, directors, and VPs. They are either technical or used to be technical, but the common denominator is that they have a million things on their plate. You are communicating the what and the why, so the core idea needs to be memorable. Historically, this might have been a detailed powerpoint or a concise Google Doc. What I have been doing lately is writing everything in markdown and then having an LLM convert it into an HTML/CSS/JS presentation. The front of each slide stays high level, but clicking it reveals a back side with underlying details. Definitely have an LLM make the architecture and sequence diagrams. If something is confusing, make the diagram interactive. LLMs are really good at this. If I really need management buy-in, I would prototype a demo that shows what the final result could look/feel like and link to it. 3. Low-Level Doc - De-Risk For the low-level doc, the audience is whoever will be doing the implementation. Historically this meant engineering teammates, but these days it means agents too. Since people can use LLMs to ask what something means, I generally keep this doc pretty technical. If part of a feature is ambiguous or hard to reverse, I will usually include a few options and do a spike on each one before deciding which direction to commit to. This doc will be full of pseudocode snippets and links to the services and files that the implementation will touch. By the end of the document, there should be a very clear answer for how we will address each of the "whys" from the high-level doc, what metrics we will use to measure the success of the rollout, and how the project will be released in stages that build on top of each other. For key architecture decisions, I usually evaluate the options across a few dimensions: 1. Scalability: What happens at 10–100x the current load? What happens if 40 engineers are contributing instead of 4? 2. Reversibility: How painful will it be to change this decision later? 3. Ownership: Who is actually going to own and operate this system? Do they have the bandwidth and track record to maintain it well? 4. Ability to reuse existing interfaces: Can we reuse something that already exists, or are we introducing a new boundary? 5. Verifiability: How easy is it to prove that this system is working correctly? Can we test it, observe it, and quickly tell when something is wrong? 6. Cost: What does this cost us now? What will it cost us as we scale? Some examples of this might be: * "Should we spin up a starrocks instance or reuse our existing postgres instance?" * "Does this existing microservice meet our TPS and latency standards, or should we build our own service for this functionality?" * "Is there an MCP an agent can use to read data from this analytics warehouse?" 4. Execute in Stages - Execute Generally I like to build towards a final solution and not just one shot it, even though it is tempting to do that with AI. I think you should still deliver something in stepping stones, or else it might turn into a skunkworks project that is always 2 weeks from being done. There is some version of the project you can ship in 2 weeks, another in 2 months, and another in 6 months. Build towards those. Every two weeks you should be at a point where you have something worth demoing. 5. Shipping Means More Than Releasing to Prod - Validate & Iterate Something isn't shipped when it is released to prod. It needs to have: * Clear verification loops * Success metrics * Observability and monitoring * Analytics * Rollout, adoption, and awareness (from users and the org) So my basic loop is: understand, align, de-risk, execute, validate, iterate.
2
6
236
someone with 10+ years of experienced would know to use "begin;" before running the command. strong disagree an experienced engineer would ever do this. this is like sql 101.
> vibe coder someone with 10+ years of experience is honestly more likely to make this mistake than an llm
28
slop means different things for different domains. slop for code means bad code. but art is subjective. even if something is technically "good", if it becomes too familiar, it won't be interesting and will come off as generic and boring aka bad. There is a reason painting eventually had to move away from realism once artists could consistently draw near perfect visual reconstructions. one problem with using models for art is they are lossy. unless you have very very specific prompts and reference material they will naturally fill in under specified details with the model's learned defaults, producing familiar outputs this was a huge problem with SD1.5. It would create this amazing looking photo of a woman, but after seeing 100 of them, they all started to look like sort of the same woman. Similar facial proportions, lighting, composition, skin treatment, depth of field, etc etc so I do think that once people are exposed to enough art made by a familiar AI model with the human largely out of the loop, even if the art is technically good, people will stop finding it interesting. x.lingyaoai.com/_AashishReddy/status/2…
Some people have the hypothesis that actually, Claude-like writing is fine-in-itself, and it only annoys people because they're over-exposed to it. I now see people praising the visual style of Opus 5.5, and many people are using it to make stuff. In the near future, will such material be thought slop? Poll below.
82
A lot of b2b ai startups still sell "make your employees more productive." But like everyone is 100x more productive than they use to be a couple years ago. That really isn't a problem that people are in need of solving right now.
24
I was wondering why the Picasso diagram (great way of visualizing join algo selection) didn't include merge join... but thinking about it... this is a distributed system so even though each shard is sorted locally, the data isn't sorted at a global level. so when you fetch the data across the shards it is not pre-sorted, a key advantage when deciding to use the merge join algo. this means you either have to do the sorting yourself or at the very least a k way merge. planetscale.com/blog/the-lif…
1
64
yeah so pstack is actually very good. funny how powerful just a bunch of markdown files can be.
45
Hot take - AI makes working with bad software engineers much more painful. Historically, bad engineers were somewhat rate-limited. If they were not very smart they couldn't write a lot of code. There was a correlation between writing code and being smart. And thankfully eng strength rate limited how much bad code they would write. Now you can be dumb as bricks and write a boat-load of code. How smart you are has nothing to do with how much code you can ship. Ironically this slows down the strong engineers from writing code themselves since now they have to spend all this time reviewing 2k lines of AI generated slop. I think I have heard the word "slop cannon" used to describe this phenonomen. Great moniker lol.
42
I wanted to dive into how LLM tool calling really works since the first time you use it, it feels like magic. If the model is just predicting token after token, how can it suddenly make an API call? The short answer is: it doesn't. The runtime does though. The runtime being the server that is executing the model. An LLM generates tokens indicating that it wants a tool called. The runtime outside the LLM interprets those tokens, runs the tool, and sends the result back to the model. Let's break it down: 1. LLMs predict the next token At the most basic level, a transformer receives a sequence of tokens: """ [x1, x2, x3, ... xn] """ And produces a probability distribution for the next token: """ P(x[n+1] | x[1:n]) """ Then it picks a token, appends it to the sequence, and does it again: """ prompt ↓ predict token ↓ append token ↓ predict token ↓ append token ↓ ... """ Nothing about that pattern says: """ HTTP Bash Code exec """ It just predicts tokens. 2. So how can predicting tokens cause an API call? Model trainers give certain outputs special meaning. huggingface.co/docs/transfor… At a high level, the model might generate something like: """ <tool_call> { "name": "get_weather", "arguments": { "city": "Tokyo" } } </tool_call> """ The exact representation depends on the model and provider. But the point I want to drive home is from the model's perspective, it just generated another sequence of tokens. But from the runtime's perspective, which has been observing the token outputs, that sequence means something - Stop generating text and return a tool request! The model did not make an HTTP request. It predicted that get_weather was the most appropriate thing to generate next. The software around the model is what gives that prediction real-world consequences. 3. Your server actually calls the tool For a custom tool, the flow is roughly: """ Your server ↓ sends prompt + tool definitions ↓ Model provider ↓ LLM generates a tool request ↓ Model provider returns the request ↓ Your server calls the weather API """ The model runtime pauses execution when it notices a tool call is desired and calls your server. At a high level, your server might receive something like: """ { "type": "function_call", "name": "get_weather", "arguments": { "city": "Tokyo" } } """ At this point your agent harness takes over: """ if tool_call.name == "get_weather": result = weather_api.get( city=tool_call.arguments["city"] ) """ The LLM proposes the operation. Your application decides how to run it and performs the actual IO. The model saying: """ delete_customer(id=123) """ does not mean the customer should be deleted. Your software decides that. Your harness most likely will have authentication, authorization, input validation, other side effects around deleting a customer, etc. 4. The tool result becomes input to another inference step Suppose the weather API returns: """ { "city": "Tokyo", "temperature": 27, "conditions": "Rain" } """ Your server sends both the model's tool call and the result to the conversation: """ { "role": "assistant", "tool_calls": [{ "type": "function", "function": { "name": "get_weather", "arguments": {"city": "Tokyo"} } }] } { "role": "tool", "name": "get_weather", "content": "Tokyo is 27°C with rain." } """ After your application runs the tool, it adds the result to the conversation. Before the next inference step, a model-specific chat template turns those structured messages back into the exact format the model expects. """ API / tool result ↓ structured conversation ↓ model-specific chat template ↓ serialized text ↓ tokenizer ↓ token IDs ↓ embedding layer ↓ vectors for the next inference step """ That content gets serialized into the model's expected message format, tokenized, embedded, and processed like the rest of its input. Then a new inference step begins and the model can now predict: """ It is currently 27°C and raining in Tokyo. """ So the complete loop is: """ tokens ↓ tool request ↓ normal software ↓ API response ↓ more tokens ↓ more inference """ What happens to batching while the tool runs? When reading this, I am sure you thought about the performance implications of this and how it might affect batching. How are model providers able to most efficiently use the gpu, given they do not want the GPU sitting around waiting for your weather API. Once the model finishes generating the tool request, that inference step is done. The provider can use the GPU capacity for other requests while your server is making the API call. When your server sends the tool result back, it becomes another inference request that can be scheduled onto the GPU. So tool calling is less like pausing a function halfway through: """ weather = await get_weather() """ And more like ending one model invocation and starting another: """ return weatherPromise """ In an inference engine such as vLLM, they might use the term batching, but under the hood it is really a complicated scheduling algorithm. vllm.ai/blog/2025-09-05-anat… After each model step, finished requests can leave and new requests can join. Your tool call ends one inference request, and the tool result returns later as another. The weather API adds a gap to your workflow, but it does not reserve an empty seat on the GPU while everyone waits. Disclosure: This post is a summary of a conversation I had with chatGPT and the sources it gave me about the subject
130
Why is it better for an LLM judge to use pairwise comparisons or a small set of enums instead of numeric scores? Numeric scoring turns judgment into a regression problem. LLMs are better at fuzzy semantic decisions than precise estimation: they have a lossy latent representation of the world, not a precisely calibrated one. In general I would say LLMs excel at: ranking > classification > regression
59
Ai definitely seems to be effecting people’s willingness to pay for software. I don’t think that it is so much as people are building their own software. But that things like Claude code and codex are becoming the everything app.
App Store sales dropped for the first time in a decade
51
VCs trying to get their instinct take out
china’s AI robot just hit 14.5 m/s
46
one nice thing about ai + agents advancement is that inbox 0 is finally possible
33
by kv cache miss there seems to be two types: user - changed system prompt, moved message order around, injected something dynamic towards the beginning, etc infra - request hit different instance, cache evicted due to memory, instance rebooted, etc
29
wow a hit tweet! us "read the code" truthers right now
Replying to @marcba
a compiler has never said "Now I see the real issue" and then been completely off
48
AI has let people develop systems that cannot be managed or reasoned about without the assistance of AI, creating a lock-in effect and a dependency on the models. Being able to utilize a complex system can be a competitive advantage, but it can also put you in a position where it is difficult to extend the system and it cannot even be used without spending money. IMO AI has made tech debt more manageable, not less. Now the type of debt an org seems to take on is knowledge debt - the fact that no one deeply understands the system. Knowledge debt is working software that no one fully understands anymore. The code runs, the product works, but the team no longer has a clear mental model of how it all fits together. It’s insane how much code is being shipped now. It used to be a struggle to keep an entire system of microservices in your head, but now a single service can be a challenge. Haven't reviewed PRs in a week? Well here are 10k lines of code for ya. Now you are at the mercy of LLM code summaries not hallucinating. There is a lock-in effect where now you are dependent on LLMs for you work for even the most basic changes. Once again though, this is debt. You can pay it down. Take a week and build the mermaid charts, HTML visuals, integration tests, and whatever else helps the team rebuild a mental model of how the system works at an engineering level, not just a product level.
36
The reason why I think bad AI writing bugs people so much - it is genuinely disrespectful to ask someone to put effort into reading something you couldn’t be bothered to put an ounce of effort into writing. AI makes writing cheaper but in some ways it makes reading even more expensive. AI writing is overly verbose. If you generate five paragraphs in seconds and just give it a quick skim, you have saved yourself time by creating work for the readers. You still need to read it and distill it. That’s the job now!! If you don't want to be a writer you still need to be an editor. AI-generated PR descriptions created solely from the code diff. They might sound professional and explain the code well, but they often miss the overarching business context. An even more egregious example is PR reviews where someone asks an LLM to review your code, pastes the response, and clearly doesn’t read it. Like I can do that lol! I also have access to the same models... What value are you adding?? The craziest part is when the AI output says something like: ``` The PR is well structured and handles X, Y, and Z effectively. A few improvements could be made to Y by doing A or B. The author is also asking for guidance on how to handle X. You'll still need to provide your own perspective on that. ``` Had they simply read the response before sending it, they would have realized the LLM was telling them they needed to do the tiniest bit of thinking themselves.
3
210
would really appreciate a follow-up of this video by @karpathy : piped.video/watch?v=7xTGNNLP… would be great to cover what pre and post training look like now, how the old jagged edges were solved (strawberry problem), what are the new jagged edges, etc.
1
63
watching a year and half old kaparthy video and noticed that the gpu he was like 25% cheaper than the exact same gpu from the exact same provider now lol
1
55
feels like the most successful neural network arch is designed 1:1 with real world hardware limitations
1
54
most obvious example is the transformer allowing for parallelism that lstm did not. but a more precise example is moe maps to gpu rack layout.
1
32
feels like it makes a lot of sense for labs to either be very involved with their hardware partners or own hardware
20