Member of the Technical Community. Captain of @myfirstmate @theSSHHIP Author of axi.md

Bellevue, WA
firstmate just crossed 3k stars on github! for most of my other projects, people would call them “helpful”. firstmate was the first project where i repeatedly heard people say it “changed their lives” - and i know exactly what they meant, because it completely changed how i work too reflecting on its journey towards this milestone, i’d like to share the story of developing this project and the key philosophy that made it special i created firstmate a few months ago, at a time when i constantly looked at a long list of parallel agent sessions, and had to juggle between them every time i jumped to a different session, i had to context switch and try to remember what the hell the session was about, what the agent was saying and what the right next step should be most agent harnesses and orchestrator apps simply made it easier to see the sessions and jump between them. but i still had to context switch all the time. i’m pretty sure there’s a KV cache in my brain and it was overheating i thought that’s a terrible experience for developing software, and it can’t possibly be the end game. i used to enjoy coding - the focus, the flow, and the peace. watching it replaced by this non-stop tab-juggling gave me more pain than watching game of thrones season 8 so when i explored firstmate as a solution, the biggest thing i wanted to accomplish was not a smart workflow, or a useful tool, or an impressive piece of technology it was to create an experience. a feeling. a sense of peacefulness that was taken away by the noisy agents. a sense of confidence that everything’s under control. an ease of mind that nothing will fall through the cracks the moment i look away it’s the experience of being a good captain that sails with a well-managed crew, with the help of a firstmate that carries out the captain’s direction i spent lots of time figuring out what it takes to create such an experience. i could build yet another agent harness, call it “open ship” or something. or i could build another orchestrator app, automux but eventually i discovered that the experience i’m trying to create is completely orthogonal to what harness or orchestrator people use - it’s a new way of working, and it doesn’t discriminate what tools people use that led to everything that firstmate is today - it’s an agent distro you can run with many existing agent harnesses like claude code, codex, pi, and across many session managers like tmux, herdr, orca it introduces the new way of working through a memory file that steers your agent to perform the role of an orchestrator, a bunch of scripts that handled deterministic logic, a set of agent hooks that improved reliability and efficiency, some built-in skills like “ahoy”, “afk”, “stow” that handle common operations, and some customization such as the “calm” mode - all designed to create the experience i had in mind from the very beginning the added benefit of being an agent distro this way is that the agent can fully self-introspect and self-modify. if something’s not working well, just ask firstmate and it’ll figure it out (now the unintended side effect is that i have 500+ PRs coming from everyone using their firstmate to self evolve). it also became very easy to setup - clone the repo, run your agent in it, that’s it that’s the story of firstmate - a pursuit of an experience leading to a solution that created quite a special feeling. if you know, you know github.com/kunchenguid/first…
62
32
435
102,043
coding is solved. where is the moat? new apple guy tweets one word “hello” and gets 90 million views there’s the moat
13
5
151
8,054
Kun Chen retweeted
I made fleetdeck, a terminal dashboard for firstmate by @kunchenguid firstmate keeps all fleet state in plain files. fleetdeck reads them and shows one screen: - tasks, PRs, CI and conflicts - open decisions and second mates - the context each session loads - activity events
1
1
13
1,526
there’s now a new mental model for the software stack - skills are the new “apps” you install them to get more functionality. people are starting to build and share skills more than real “apps” many skills are very simple today but i think we’ll eventually see increasing sophistication - mcp servers and CLIs are the new “APIs” they allow services to encapsulate their core capabilities behind an interface the can be invoked in a composable way many skills will be calling mcp servers or clis behind the scenes just like how a lot of apps call APIs today - agent harnesses are the new “operating systems” it sits between the user and the underlying resources, helping the user manage common concerns such as compute, memory, security etc if you look closely, it’s hard to unsee the resemblance - claude is the new macos. codex is the new windows. pi is the new linux skills run in agent harnesses just like apps run in existing OS - LLMs are the new “computers” they process instructions coming from the user, the skills, the agent harnesses, and compute the output the model weights is the cpu. context window is the memory. actual computers are becoming more like the power source ok so why do we need this mental model? i think it helps us think about where we play if you want to be the new app developers - start to learn to write good skills if you want to build a saas, figure out how to package things as good mcp servers and CLIs if you want to be linus and write the new linux, build an agent harness. otherwise, you can still try to learn about how agent harness work like how CS students used to learn about OS if you want to be the new apple or samsung, start to learn how LLMs are trained and play with open weight models - you may end up inventing the new raspberry pi!
58
39
408
14,882
people who laugh and say karpathy's latest tips are outdated - i suspect the vast majority of them have not even tried the tips yet i just tested the ASD-STE100 wording rule and it's surprisingly good at helping increase clarity of model response, even within html artifacts. but the trick is that the full ruleset is a bit too strict and you need to pick a subset one prompt you can run super easily: "randomly sample 10 session transcripts where i worked with you interactively within the past week. apply ASD-STE100 rules to assistant responses and analyze which rules would have increased clarity, reduced confusion and improved the conversations, then document those rules in my user level AGENTS.md" you might be impressed!
We'll be spending a lot more time trying to understand the outputs of language models. A few thoughts, tips & tricks: Writing. Something I've had success with: Ask your LLM to explain something in ASD-STE100, it's a controlled language specification originally developed for aerospace maintenance documentation. LLMs well-versed in this language and it comes with heavy constraints on clean writing style that I often find a lot more readable. Sometimes I've tried to soften it a bit e.g. ask for "80% of the way to ASD-STE100" because the spec is quite stringent. But even better: Diagrams / images. Instead of writing, ask your LLM to create a diagram. These can be a lot easier to process, parse, and understand. But even better: Web pages. Ask for output "in HTML" to get a beautiful, interactive webpage. LLMs are getting really good at frontend and can create beautiful experiences, animations, etc. But even better: Explainer videos. The output format I am most bullish on is fully custom / bespoke explainer videos generated on any arbitrary topic. Experiment with things like "Create a 3b1b style video explainer on X. Use my ElevenLabs API key for audio narration". (you'd need an API key for the latter or you can ask your LLM to find you decent free alternatives that use your local compute). This is actually starting to work! In summary: - As LLMs get better, they will do more and more of the legwork autonomously, and a lot more of our work will rise up the abstractions into oversight and understanding. - Luckily, LLMs can help here too because as intelligence and code are increasingly abundant, you can ask for large, custom, discardable software artifacts (e.g. web apps, video explainers) that would have never made sense to create before. Push the boundaries here and you'll be surprised.
81
192
2,962
270,447
Kun Chen retweeted
Put Kun the @grok @bot to use right away. Started a PR last night on a live booking agent. Four SMS booking gaps turned into: 28 commits in ~11 hours... review>fix>re-review with cloud rounds green and GitHub red. It reviewed the whole PR and said stop the loop and freeze the bar immediately. Work needs to freeze when the original done bar is met, or when the next fix is a new requirement. We locked that into a standing skill for the @bot crew. Name: PR freeze bar Use when shipping a PR, running adversarial review, or deciding whether a finding stays in the current PR vs a follow up - write the done bar before work starts and freeze when it is met. Gist: gist.github.com/SSBrouhard/5… Already paying off 🫡
Kun the grok @bot is here! 😀 principal engineer at your service - x.ai/bot/xK8W0ukRv4iZjglzz-F… this bot is continuously updated - everything i posted, shared in a video, or open sourced on github - all available through this bot hope i can be helpful to more ppl this way!
4
1
4
1,924
sigh.. must say a few things about this is face to face video chat the most efficient way to communicate with a machine? imagine you want to cancel your netflix subscription. would you prefer a button click, or turn on your camera and talk to a bot face to face? how about booking a hotel or flight? excited to do it via a video chat with a customer service bot? imagine spending all day in zoom meetings for work, and now it’s all zoom meetings after work too? face to face communication is great at perceiving human emotions and feelings so humans can get on the same page ands build rapport with each other - this makes zero sense when it’s with a machine, except for things like therapy now let’s think harder - if such tech is widely available and actually good, which industries have the most incentive to adopt it? well, porn and fraud come to mind imagine people stopped dating real people because “AI girlfriends” are always available imagine grandmas getting a facetime from their children imagine job interviews are just one bot talking to another, and the more expensive model wins this is a technology that seems far more easily to get used for harmful purposes than for real economic benefit i hope people tread carefully here. build it because we know it can do good, not just because it looks cool as a demo
Introducing Griffin, the first model to pass the video Turing test. 48% of people who talked to it live thought it was a real human. Previous systems have had a pass rate <3%. It is #1 on NVIDIA's benchmark for full-duplex AI video. It’s the first Human Interaction Model (HIM).
Community note
The 48% figure and "video Turing test" claim are from Tavus's own study of 54 one-minute calls, not independently verified or using a standard protocol. Griffin-Lite leads NVIDIA's VideoFDB benchmark on their public leaderboard. cellcog.ai/blog/tavus-gri… research.nvidia.com/labs/amri/proj… tech-ish.com/2026/10/02/tav…
66
16
283
24,158
the core reason i’m nervous about this is that fundamentally, this is about manufacturing human emotions just think - what can a fake human in video chat do, that today’s voice agent can’t? it’s the facial expressions, body language etc that represent and create emotions when we start manufacturing that at scale, we are removing a very precious thing from humanity that’s why this needs a lot more thoughtfulness than what’s demonstrated here
3
3
44
2,006
Kun the grok @bot is here! 😀 principal engineer at your service - x.ai/bot/xK8W0ukRv4iZjglzz-F… this bot is continuously updated - everything i posted, shared in a video, or open sourced on github - all available through this bot hope i can be helpful to more ppl this way!
13
13
148
10,939
some people dunk on the bumpy demos on openai devday i however really loved how everything was actually real, how the hiccups were handled with composure, and how they came back with this real humans, real work, real emotions - really earns my respect
dots demo. Now... with better WiFi.
38
11
456
26,334
if you like drama, here’s some drama for you a few grown men arguing in public matan: i dumped chris chris: no no no i dumped him FIRST! scott: yep chris dumped matan first, for me matan: want to see some emails?
We are terminating Chris Degnan for unethical conduct involving Cognition. The last few months have seen incredible progress in AI capabilities. San Francisco has flourished as new companies that solve new, more ambitious problems are finding great success. Generally, it is a wonderful time to be building. We at @FactoryAI have seen overwhelming interest in our model-agnostic coding agents and our software factory product. Our team has 10x’d in size, while our revenue has 100x’d year over year. This momentum has been unprecedented. A much larger competitor, Cognition (makers of Devin) has fallen behind us on the capabilities that matter most to customers: cost, quality, and security. Instead of competing in the market, Cognition engineers feigned interviews with us to pry information about our product. Not finding what they were looking for, Cognition has decided to throw their weight and money at people with direct knowledge of our most confidential plans. It feels as though ethics is being thrown out of the window in the AI era. People are willing to do anything, including exploiting privileged information and violating ethical boundaries. I think this is unique to our time, and I don’t think it’s right. Integrity still matters. Yesterday, I made the decision to immediately terminate Christopher Degnan’s roles as a Board Observer and Advisor to Factory, after over a year of service. Prior to this decision, Chris told me he had a casual conversation with an executive at Cognition AI. When I questioned his intentions, he reassured me that ethics aside, he had “made too much money” and was “too lazy to go work for Cognition,” which I trusted and believed. On Monday, Chris spent time advising the Factory team on a handful of confidential board-level matters. That evening, he called me to say that the conversation that was initially described as casual and one-off was actually formal and recurring. For weeks, while he sat in our board meetings and advised our leadership team, he was also confiding with executives of our largest competitor. Chris was subject to confidentiality obligations in connection with his work with Factory. We do not know the extent of the information he shared, but it puts his timely questions about our product roadmap and what the parity gap involves into a new light. It is sad to sever a relationship with someone who has been a trusted advisor - and even a close friend - for over a year. But Chris’s conduct is unacceptable to me. Trust in Board Membership is one of the sacred bonds in the Silicon Valley, one that helps the startup ecosystem thrive. With it comes an enormous responsibility. That trust was violated. Competition is good. I respect and in many cases admire our competitors. San Francisco is a beautiful, singular place where the bold and ambitious go to defy the norms and precedents of the past. But certain principles that must remain. Violating ethics to seek advantage turns what should be positive-sum into zero-sum. We all love technology. And we all love to compete. But we must hold ourselves to a higher standard. The future of software engineering comes with abundance that will impact every person on Earth. Building that future comes with immense responsibility. We must build and compete with integrity.
9
3
86
21,142
it's mind-boggling google would publicly announce something the public cannot use like.. why would someone scream "I'M NOT BEHIND" from behind? but then i just realized - October is about time people in google start writing their performance review packets...
Introducing Gemini 4 Argon – our new frontier model. It’s built for complex workflows across coding, enterprise knowledge work, and cybersecurity defense – rolling out today to a set of trusted testers through our Fairwind Program.
110
43
1,361
128,994
used sol 6.1 for a whole day and i'm very.. confused will try my best to break down the key things to know 1. sol 6.1 reminds me of opus 5 this is a model that makes me constantly have to ask "wtf are you talking about" it's very capable. most of the time once i understood what it meant, it did mean the right thing. but it's a pain to get through its choice of words i hope there's a sol 6.5 coming that ends up being what opus 5.5 was to opus 5 and get this fixed 2. it's bigger than previous sol you can easily observe the slowness that comes with this model. this hints at it being a bigger model than 5.6 sol. but my confusion is why it's priced at terra and sonnet range then i realized, the pricing is set to attract enterprise customers who pay at API rate for consumers who buy the $200 subscription, the lower price doesn't mean more usage because they're cutting the subscription's quota by half so i guess the intention is to give enterprise customers an opus competitor at lower price, while holding back the benefit from consumers 3. all gpt models still suffer from poor judgment sol 6.1 is likely taught by astra. if you read some of its code, it smell like astra code. it also inherited astra's strength which is that it can get stuff done with less turns and tokens (even opus 5.5 hasn't got to this level of efficiency yet) but... the biggest problem to me is still "judgment". this is most obvious when i start to use either astra or sol 6.1 as my firstmate - everything simply starts to fall apart because of poor decision making for example it would see me approving a plan and decide to file it in the backlog rather than the more obvious next step which is to start implementation (which opus 5.5 and fable would always "just know" it should do) all signs point to a lack of human judgment in post training - all gpt models are likely heavily optimized by RLVR overall - i'd categorize sol 6.1 as a solid implementer and reviewer, but i would not use it interactively openai has work to do - right now they simply don't have an opus 5.5 equivalent anywhere in its line up, which is problematic but seeing how anthropic could come back from opus 5, there's still hope :)
87
24
795
85,072
comparing dots, muse, and grok bot just made me realize something important grok bot is the ONLY one that made sharing bots easy i came across quite some grok bot users in the wild using the Firstmate bot i shared x.ai/bot/__4FfrkUdvpdMk6-LKg… (this is basically how i use grok bot btw!) i can see openai may be hesitant to do something like this because "custom GPTs" never quite took off (anyone still remember it?) but i think this time it's a bit different. these cloud agents now all have their VM, a code execution environment, and a ton of connectors - they can do far more things than what custom GPTs could do when it was introduced with this explosion of possibilities, prepackaged bots now make more sense because people want proven playbooks and not having to reinvent the wheel each time a marketplace like this also creates some network effect which is good for building an ecosystem moat. this is something i think dots and muse should both think a bit more about
17
8
185
12,894
why is real world applications of robotics still so behind? literally none of my vacuum cleaner robots ever finished a single job without getting stuck one way or another where did all the research go??
31
4
123
14,280
this is my dot. his name is - surprise surprise - Firstmate!
20
3
203
11,235
something to internalize about ultrafast mode 8x speed at 6x cost means it will cost 48x $ per second it might be okay if it got the job done. but if it ever went down a wrong direction, it will be burning astra $ at 48x speed before we even have time to see and stop it
23
4
179
16,299
ok here's my key takeaways from OpenAI DevDay keynote 1. dots grok bot, muse, and now dots - they may look similar, but they are not i think openai's dots is a significantly more cohesive implementation that spans across personal and business use, and also across single player and multiplayer pretty excited to see where it goes 2. sol 6.1 biggest thing? openai caught up on cheaper cached tokens. this is a big deal and affects the overall cost of long horizon work a lot in terms of sol the model itself, honestly it's bit weird that 6.1 is released only a few days after 6, making it feel like sol 6 was a premature release and 6.1 is what's supposed to be sol 6 all along we'll just have to put our hands on it to see how big of a jump it is 3. jev clone on one hand, it's incredible to see how quickly they cloned it. but on the other hand, it's scary to see how these frontier labs would not blink an eye to destroy smaller startups in their space openai's decision model already has image understanding which jev doesn't. it speaks to the moat openai has accumulated which makes it hard for startups to disrupt 4. ultrafast 8x speed at 6x price, only on astra right now but said to be coming for sol as well i don't know... how many use cases can actually get a positive ROI using astra at 6x price? 5. $500 subscription with the adjustment to the $200 plan and the introduction of $500, it's become clear that subsidization is coming to an end $200 is simply 10x of what $20 gives, and $500 is simply 25x - zero discount at these higher price points i'd rather they create $100 "packs" that you can stack onto your account 6. sign-in with chatgpt, and marketplace this is a BIG DEAL!!! why? because it means other companies can finally build agentic products without having to sell tokens selling tokens is a very bad business model for the vast majority of SaaS companies. it complicates their pricing structure and creates almost no profit margin at all this move from openai means customers just buy tokens from openai, and use them everywhere - very good for the ecosystem the impact to the broader industry will be amplified if anthropic gets pressured to follow the same pattern alright - might share more thoughts after actually putting my hands on the new things, but those are the initial reactions for now. curious to hear what you all took away from it as well!
74
91
1,439
167,390
so @sama still calls X "twitter"
47
122
11,709
oh no - codex $200 subscription's dollar value is being cut in half in their defense, openai's models are indeed quite efficient. but anthropic has caught up too and the big cut is only the first order effect the second order effect is that this will give anthropic enough room to cut their subscription value as well the third order effect is that the labs have found a way to gradually end the subsidization for consumers subscriptions. just do this a few more times, and we'll see the subscriptions charging the same as API pricing i predict that in about a year, we'll refer to what we had as the good old days
Hi, Tomorrow we are re-opening the Pro $200 subscriptions to new subscribers, but together with it we are also changing how we calculate the usage for it. In effect, if you do the math, it will net out at half the dollar in API spend compared to the old Pro $200 plan. Now that it's said, let me explain why this is happening and why you will still get more work done than if you were on the Pro $200 subscription one month ago. (a) We didn't want to compromise in other ways and are committing to not reintroducing the 5h limit, so that you can fully use the weekly usage when you want. (b) On the subscription, we guarantee that over time you always get more work done and with an increasing level of quality. This means that you will continue to get more value per dollar spent as a result of models getting more efficient and us passing down the improvements in the form of API price reductions. (c) We don't want to put an incentive on ourselves to artificially inflate the API list prices to make it look like you are getting a lot (and workaround it through discounts, etc). Instead we want to continue to both rapidly reduce prices and increase capabilities of models on the API. This week we introduced GPT-6 Sol and GPT-6 Luna at 50% of their previous price. Over time, we see prices go low enough that it makes sense for most to buy usage as needed without there being a significant gap between what you get in a subscription and what you get in the API for a dollar spent. (d) Tomorrow, we are adding more things to the subscription that won't draw on the usage, I won't reveal what that is yet. I wanted to be transparent before all the big announcements tomorrow. Lots of new exciting things are coming to the subscriptions that will make it super compelling, but I wanted to make sure to share this change ahead of time so you can all understand it before we shower you with good news. Codexingly, Tibo
140
27
905
109,645