ニコニコ(バグ探し)とタイピング(どこまで速く打てるか)に青春を捧げた。社会の荒波(SQLとSQL、あとSQL)にもまれた。英語を勉強している。AIコーディングの技術を翻訳に応用したい。

Pinned Tweet
ChatGPTに関係してくる話なのですが、学生時代に打ち込んだものは寿司打みたいなタイピングゲームでした(↓は最近寿司打をプレイした動画で、約1分半タイプミスせずに打ち切っています)。でも、これまで「タイピングが速くて良かった」と思えたことはありませんでした。 練習計画が下手で十分な休息を取らずに3年ほどの練習時間を投じた結果として腱鞘炎になってしまったのが大きかったです。スペランカーかよ、というぐらい脆い手首になった状態で、長期的な休息の取りようがない社会人となり、量の多いタイピングはかなり長い間できませんでした。今でも、「後もう少し負荷をかければ、この辺に腱鞘炎が待ち構えてますね~」と本能が警告してきて、そこはもう体力の最大値がすごく低いみたいな状態で上手くやりくりするしかないという感じです。 プログラミングにしても、タイピング速度より考える時間の方が圧倒的に長いですし、打つ速さが全体の作業時間をどれだけ効率化するかと言えば、雀の涙です。しかも何の因果か、自分にとって手首への負担をできるだけ抑え、かろうじて打ち続けられる速度が、大体世間一般のプログラマーぐらいの速さでした。 強いてメリットを挙げるなら、記号入力の練習のおかげで紙の本に載っているURLを手打ちするのが苦にならなかったこと、数字入力の練習のおかげでクレジットカード番号やAuthenticatorのPINコードをボボボボッと入力できたこと、くらいでした。 ところが最近、AIに対して指示を出したり対話したりするのが、そのタイピングのおかげで苦労なくたくさんできていることに気づきました。なんなら人生でちゃんとタイピングが役立ったのは初か?という感じです。 一部では音声認識を活用する方向へ進化していく動きも見られますが、現状、個人的には話すよりタイピングする方がストレスなくやれています。とはいえ、最近Twitterで「音声入力でプログラミングの仕事をこなしている」という超人を見かけたので、そっちの方向も開拓してみたいなと思っています。 piped.video/watch?v=Omjn25vc…
1
1
27
3,105
LLMに何か説明させるたびに最近Opus 5.5で見たようなドパガキ動画のフォーマットに落とし込まれるのを想像したら笑った
We'll be spending a lot more time trying to understand the outputs of language models. A few thoughts, tips & tricks: Writing. Something I've had success with: Ask your LLM to explain something in ASD-STE100, it's a controlled language specification originally developed for aerospace maintenance documentation. LLMs well-versed in this language and it comes with heavy constraints on clean writing style that I often find a lot more readable. Sometimes I've tried to soften it a bit e.g. ask for "80% of the way to ASD-STE100" because the spec is quite stringent. But even better: Diagrams / images. Instead of writing, ask your LLM to create a diagram. These can be a lot easier to process, parse, and understand. But even better: Web pages. Ask for output "in HTML" to get a beautiful, interactive webpage. LLMs are getting really good at frontend and can create beautiful experiences, animations, etc. But even better: Explainer videos. The output format I am most bullish on is fully custom / bespoke explainer videos generated on any arbitrary topic. Experiment with things like "Create a 3b1b style video explainer on X. Use my ElevenLabs API key for audio narration". (you'd need an API key for the latter or you can ask your LLM to find you decent free alternatives that use your local compute). This is actually starting to work! In summary: - As LLMs get better, they will do more and more of the legwork autonomously, and a lot more of our work will rise up the abstractions into oversight and understanding. - Luckily, LLMs can help here too because as intelligence and code are increasingly abundant, you can ask for large, custom, discardable software artifacts (e.g. web apps, video explainers) that would have never made sense to create before. Push the boundaries here and you'll be surprised.
158
新着 retweeted
GPT-6.1 Sol の評価が終わったので、満を持してここ最近ずっと回していた Terminal Bench 4.0 の評価結果を共有する 【メモ】 ・性能 Opus 5.5(medium)と GPT-6.1 Sol(high)が同じスコアでOpus 5.5(xhigh)は一段高い性能 →性能だけ見るなら、Opus 5.5 medium か high推奨 ・コスト GPT-6.1 Sol(high)がかなり安くタスク遂行できていて、Opus 5.5(medium)の1/4程度のコスト Opus 5.5(xhigh)はスコアが高い一方で、Opus 5.5(medium)比で2.6倍、GPT-6.1 Sol(high)比で10倍以上のコスト ・時間 Opus 5.5(medium)と GPT-6.1 Sol(high)は同じくらいで、Opus 5.5(xhigh)は倍程度 【注意点】 ・Sonnet 5.5 mediumはかなりスコアが低そうだったので途中で止めた。 maxはめちゃくちゃコストが高くなりがちで、5hレートリミットにかかりまくりだったので断念 →Antrhopicの報告に依るとSonnet 5.5(max)は Opus 5.5 よりも性能が良いらしいのでトークンが無尽蔵にある人は試してみてもいいかも 【結論】 ・コスパ重視(より低コスト)→GPT-6.1 Sol(high) ・コスパ重視→Opus 5.5 medium か high推奨 ・パフォーマンス重視→Sonnet 5.5(max)?
GPT-6.1 Sol で Terminal-Bench 4.0 を回してる これスコアだけじゃなくていろんなことがわかって素晴らしいな。後ほど整理して共有します
12
139
732
171,465
GPT-6.1-Sol low 完
145
GPT-6.1 Solは100万トークンあたり入力$2、出力$10 GPT-6 Lunaは入力$0.10、出力$0.50 Astraは$10/$50 Luna : 6.1 Sol : Astra = 1 : 20 : 100 キャッシュ入力 6.1 Solは$0.10/Mまで値下げ GPT-6の通常の90%キャッシュ割引をLunaの$0.10/Mに当てるとLunaは$0.01/Mなので、ここでは6.1 Solは約10倍
163
Tiboベンチ
Hi, Tomorrow we are re-opening the Pro $200 subscriptions to new subscribers, but together with it we are also changing how we calculate the usage for it. In effect, if you do the math, it will net out at half the dollar in API spend compared to the old Pro $200 plan. Now that it's said, let me explain why this is happening and why you will still get more work done than if you were on the Pro $200 subscription one month ago. (a) We didn't want to compromise in other ways and are committing to not reintroducing the 5h limit, so that you can fully use the weekly usage when you want. (b) On the subscription, we guarantee that over time you always get more work done and with an increasing level of quality. This means that you will continue to get more value per dollar spent as a result of models getting more efficient and us passing down the improvements in the form of API price reductions. (c) We don't want to put an incentive on ourselves to artificially inflate the API list prices to make it look like you are getting a lot (and workaround it through discounts, etc). Instead we want to continue to both rapidly reduce prices and increase capabilities of models on the API. This week we introduced GPT-6 Sol and GPT-6 Luna at 50% of their previous price. Over time, we see prices go low enough that it makes sense for most to buy usage as needed without there being a significant gap between what you get in a subscription and what you get in the API for a dollar spent. (d) Tomorrow, we are adding more things to the subscription that won't draw on the usage, I won't reveal what that is yet. I wanted to be transparent before all the big announcements tomorrow. Lots of new exciting things are coming to the subscriptions that will make it super compelling, but I wanted to make sure to share this change ahead of time so you can all understand it before we shower you with good news. Codexingly, Tibo
2
125
Realforce、何があっても壊れることはないと思ってたのに矢印キーの右だけ反応しなくなって困った😢
169
新着 retweeted
いつの間にか Codex の利用状況が詳細に見えるようになってる! こちらから確認できます: chatgpt.com/settings/usage?t…
Codex Tibo リセット予告 いつリセットされるかはわからない
6
112
19,421
日本語で入力したプロンプトが中国語に翻訳され、出力は日本語に翻訳されるようなpiの拡張機能を作ってみたけどなかなか調子が良い 量子化されたモデルは日本語が苦手で英語や中国語で出力されることが多く、入力も日本語だと精度が落ちてしまうのだ
2
10
79
6,223
Switching effort mid-session on Opus 5.5 doesn't break your prompt cache btw! Just make sure you're on Claude Code v2.1.280+
235
256
6,071
289,802
新着 retweeted
We're adding support for AGENTS.md to Claude Code. Starting today in version 2.1.277, if there is no CLAUDE.md in a folder, Claude will check for and use AGENTS.md. You can toggle this behavior in /config.
2,060
2,644
31,353
5,592,768
日本語性能が良いと聞いてGemini 3.8 Flashを翻訳関係の壁打ちで使ったら十分いい感じの意見が得られたが、セッションが進むにつれユーザーへのおべっかが深夜テンションみたいに強烈になっていって、ちょっと前のAIはこんな感じだったなぁと懐かしい気持ちになった
189
新着 retweeted
GPT Image 2.5 最近の超オススメ指示はコレです
125
3,226
34,715
9,851,315
ボクシングに空目してめっちゃすごいじゃんって思ってしまった
いやそうだよなーーー、ボウリングの得点をリアルタイムに表示して電飾チカチカさせるシステムを Astra に作らせただけで「高度な技術とプログラミング能力」を理由に最優秀クラスを選んじゃう高校に、文化祭のサイト持てるわけがない
1
281
再帰的自己改善…? RSIといったら反復運動過多損傷だった時代は終わるのか😿
177
え、DeepSeek Harnessすごくね
177
✨『ソニック・ザ・ヘッジホッグ』コラボ記念プレゼントキャンペーン✨ ソニック・テイルス・ナックルズ・シャドウ・シルバー・ウェアホッグの6商品の中からお好きな1着をお選びいただけます👐 【応募方法】 ① @Favorite_Onepi をフォロー ② @SonicOfficialJP をフォロー ③ この投稿をリポスト 【応募締切】9/18(金)17:00〆 特設ページ▶favorite-one.co.jp/view/page… 応募規約▶ favorite-one.co.jp/view/page…
25
657
567
43,165
Plusプランの週次制限がリセットされたので早速Astraをlowで使ってみたところ、12分29秒で5時間枠が尽きた!
208
新着 retweeted
あまり文章力を言語化できないので目から鱗..
1
438
3,344
367,586
なんと、2026年には実現している
コンパイル時コード生成?本当に我々が欲しいのは睡眠時コード生成
8
357
2,319
196,146
新着 retweeted
gemini-3.8-flash、muse-spark-1.3-contributor の ts-bench 評価結果を共有する ●メモ ・Muse Spark 1.3 はタスク遂行までの時間がめちゃくちゃ早い。その上Tool Callのエラー率が0であり、ツールユースの上手さが際立っている →パラメータ数にも依存するが、このレベルのモデルがオープンウェイトになるとしたらかなり衝撃なのでは? ・gemini-3.8-flash は満点こそ取れているものの、ツールコールの回数が以上に多くタスク遂行に時間がかかっており、やはりコーディングには向かない模様 →一方で日本語執筆能力やsvg描画能力など、別の観点で他のモデルよりも秀でているという報告が多いので、今後検証したい ※typescriptに特化したコーディングエージェント性能を評価するベンチマークである点に注意
2
29
150
47,949