🗺 Paderborn | 🚀 Investor @ Tigarius Ventures GmbH / 🌊 CEO @ WellBlue GmbH

Paderborn
Developers think like architects: coherence, dependencies, edge cases, completion. AI companies often behave like conductors: many products, many experiments, many moving parts. When launch velocity beats closure, docs lag, bugs linger, and trust erodes.
65
If you tell your Dot, “Check xyz once a day,” it creates a scheduled task in Work. However, it apparently cannot edit the model used for this. Is GPT-6-Astra or GPT-6.1-Sol used in this case?
1
44
“For Codex with ChatGPT sign-in, choose gpt-6-sol (GPT-6 Sol) if your plan and workspace provide access.” — but there's no information about the default model. - learn.chatgpt.com/docs/autom…
1
10
I heard that one in ten school lessons in France is cancelled due to teacher shortages. Imagine if someone invented AI and modernized education.
39
How do you get rid of stupid questions in your news feed, like: “If C++ is so much faster, why do AI models use Python?” or “What's stopping people from vibe coding their own AI model?” - That can't be serious.
35
The DeepSWE leaderboard was last updated on September 3. Many models have been released since then. Will we have an updated leaderboard again soon?
launching benches will never be setting a definitive and perfect goalpost in stone, rather directionally move models towards the right capabilities. we appreciate the community feedback and care to drive this together As an iterative process, we’ve been cooking behind the scenes on existing issues and new innovations to keep pushing the frontier of measurement
1
50
The silence is probably related to this:
Epoch AI just marked DeepSWE v1.1 “Flawed”: 23/113 tasks (20.3%) showed false negatives. In 18/23 cases (78%), agents editing or adding tests could break the grader - even though they weren’t told not to. Strong benchmark concept, shaky scoring.
19
AI could be the force multiplier for the next wave of cancer medicine: in mRNA vaccines, it can rank each tumor’s best neoantigens; in CRISPR, it can design safer guides/editors and predict off-targets. Both fields already use ML—better models could reduce trial-and-error.
1
32
Dominik Bödger retweeted
Global reset landing tomorrow 10am PST for all paid ChatGPT accounts. Apologies for the slow start with GPT-6.1 Sol, it's now back to running at expected speeds after the massive load spike in the first two days.
3,936
1,137
20,127
4,361,947
Dominik Bödger retweeted
This week Claude Sonnet 5.5, GPT-6.1 Sol and Gemini 4 Argon all launched near the top of the Coding Agent Index leaderboard, but each has a different balance of performance and cost The Artificial Analysis Coding Agent Index measures agents (a combination of model and harness) across three agentic coding evaluations. ➤ Claude Sonnet 5.5 (max) in Claude Code takes the top spot at 68, but also has the highest measured cost per task: $14.19 ➤ Gemini 4 Argon (high) in Antigravity CLI scores 64 at $5.84 per task - less than half of Sonnet 5.5’s cost. Note, this uses Google’s promotional pricing, and Argon is not yet publicly available ➤ GPT-6.1 Sol (xhigh) in Codex scores 63 at $1.04, roughly one sixth of Argon’s cost
94
59
998
93,630
What people always forget when talking about AI companies
Replying to @buildwithrajath
You need to look at the big picture. OpenAI's revenue has recently grown faster than its competitors, but an IPO requires solid metrics, not just billions in losses. Anthropic will likely gain more traction now, but the situation could change again in two weeks.
36
At first, I was skeptical about what this is needed for and what the benefit is compared to the remote connection from Codex or ChatGPT Work. But having a proactive assistant with a memory like this really cannot be underestimated.
dots demo. Now... with better WiFi.
26
Dominik Bödger retweeted
Congratulations to Google for achieving an excellent Canonical AA-Omniscience Index score (42) on Artificial Analysis. It did so by driving down hallucinations rather than cranking up accuracy. In other words, it probably abstained more than Opus, Fable, or GPT
2
1
11
1,028
I asked my ChatGPT dot to inspect its cloud workspace: • Debian 13.6, Linux 6.18.44 (x86_64) • 9 visible logical CPUs, AMD EPYC 9V74 • 10.45 GB visible RAM • No swap These are workspace specs, not the hardware running the AI model.
1
41
Business Premium also offers Dots in the EU. It's a great help in everyday life, even if you don't grant access to your email. The proactive notifications have been on point so far.
Is ChatGPT Dots now available in the EU as well?
1
1
43
Germany is already near the front of the AI-agent shift: 49% of office workers use AI agents. Among IT decision-makers it’s 64%, leaders 63%. Employer-provided AI use jumped 32%→49% in a year. Yet 43% of agent users fear replacement vs. 27% of GenAI-only users.
Internationaler Vergleich: Deutschland gehört zu Spitzengruppe bei Einsatz von KI-Agenten www.​n-tv.​de/wirtschaft/Deutschland-gehoert-zu-Spitzengruppe-bei-Einsatz-von-KI-Agenten-id31367738.​html
16
Working with GPT-6.1 Sol on our newsletter system: it silently changed the DOI confirmation link validity to 48h, even though the task was about a completely different part of the code. A good reminder that code review still matters even with top-tier models.
118
Dominik Bödger retweeted
woah, the biggest upgrade in gemini 4 argon will be hallucination reduction compared to 3.8 flash, google has basically crushed one of gemini’s biggest weaknesses hallucinations now close to being eliminated
gemini 4 argon scores 53 on the Artificial Analysis Intelligence Index tied with gpt-6 astra and fable 5.1, while 1 point ahead of gpt-6.1 sol
27
89
1,173
71,017
Dominik Bödger retweeted
Big news: Gemini 4 Argon (High) by @GoogleDeepMind just landed #1 in Text Arena with 1525 pts, and #8 in Code Arena: WebDev with 1679 pts! This release has reshaped the Text Arena Pareto frontier with a blended $8/MToken! Gemini 4 Argon (High) is now the most cost efficient model, see its placement on Pareto frontier below. In the Text Arena, Gemini 4 Argon (High) ranks #1 in Coding, Hard Prompts, Instruction Following, Longer Query, and Creative Writing. It also leads every occupational domain evaluated, with additional #1 spots in English, Non-English, Chinese, and Russian. This model is +20 points above the #2 ranked Claude Opus 4.6 (High), and a huge leap from Google’s previous release, Gemini 3.8 Flash (High) at #11! In Code Arena: WebDev, Gemini 4 Argon (High) gained +96 points from Gemini 3.8 Flash (High), and went from #29 to #8. Congrats to the @GoogleDeepMind team on this impressive frontier release!
Introducing Gemini 4 Argon – our new frontier model. It’s built for complex workflows across coding, enterprise knowledge work, and cybersecurity defense – rolling out today to a set of trusted testers through our Fairwind Program.
132
320
3,510
768,114
Dominik Bödger retweeted
Introducing Gemini 4 Argon, our new frontier model, rolling out to cyber defenders starting today, and more widely as soon as possible. I am really excited by the progress we have made here. Argon is priced at $2 in and $10 out during introductory pricing!
1,015
1,055
14,771
1,547,056