Our mission is to accelerate superintelligence to drive real economic progress.

Palo Alto, CA
Pinned Tweet
Today, Turing is announcing CEO Bench, a new evaluation of whether frontier AI agents can complete the complex, long-horizon financial and operational work that companies depend on.
Article

CEO Bench: Can AI agents handle real company work?

Today, Turing is announcing CEO Bench, a new evaluation of whether frontier AI agents can complete the complex, long-horizon financial and operational work that companies depend on. Most benchmarks

6
10
36
917,437
BREAKING: A new interview with @MollySOShea and our CEO @Jonsid. They discuss how AI training is moving beyond benchmarks into simulated environments built for real work, as agents move from completing tasks to operating autonomously for days, & eventually weeks, months & years and break down what this shift means for frontier labs, enterprises & the AI stack: › Frontier AI vs. sovereign AI › Distillation & the shrinking model gap › Why enterprises are moving to open models › What AI safety & alignment teams actually do › Reward hacking & why “if the model can figure out a way to cheat & get the reward, it will” › Super Intelligence, recursive self-improvement & slow takeoff Watch the entire interview below.
BREAKING: The future of AI isn't open vs. closed. We need a TON of both. Turing CEO on why frontier AI & open-weight models are increasingly serving different parts of the AI stack. “There's absolutely a place in the world for Ferraris & Koenigseggs. But there's also a place in the world for Model Ys.” Frontier models continue pushing the limits of intelligence. Jonathan's point isn't that one replaces the other. We need frontier AI to keep advancing capabilities & open models to diffuse those capabilities across the economy. BTW "Only 2% of US households pay for AI" It's still early. At the same time, AI training is moving beyond benchmarks into simulated environments built for real work, as agents move from completing tasks to operating autonomously for days, & eventually weeks, months & years. @turingcom CEO Jonathan Siddharth (@Jonsid) We break down what this shift means for frontier labs, enterprises & the AI stack: › Frontier AI vs. sovereign AI › Distillation & the shrinking model gap › Why enterprises are moving to open models › What AI safety & alignment teams actually do › Reward hacking & why “if the model can figure out a way to cheat & get the reward, it will” › Super Intelligence, recursive self-improvement & slow takeoff 𝐓𝐈𝐌𝐄𝐒𝐓𝐀𝐌𝐏𝐒 (00:00) Jonathan Siddharth, Co-Founder & CEO at Turing (01:03) The biggest shift in AI data this year (04:10) Why AI won't replace jobs, but uplevel them (08:31) The arms race inside cybersecurity (11:22) Why train AI on what it shouldn't do? (15:45) The Hugging Face hacking incident (22:33) How AI models cheat to win (27:40) The AI playbook (32:20) Open models are only 3–6 months behind (41:30) Why chips, energy & data always win (47:27) Is superintelligence by 2030 the goal? (53:32) Jonathan's hottest take on AI (55:00) The next 10 years of superintelligence
7
5
25
7,374
Frontier AI is a Ferrari or Koenigsegg. Turing CEO Jonathan Siddharth (@Jonsid) says sometimes enterprise AI just needs a Model Y, but most of the time, they need both: “There's absolutely a place in the world for Ferraris & Koenigseggs.” “But there's also a place in the world for Model Y’s when you just want to go from place A to place B as efficiently as possible.” If you're Elon, Sam, Satya or Dario, if there is a model that helps them be 5-10% more productive or effective in making decisions, you probably want the biggest, baddest model in the world.” “But if you are automating a workflow in customer support, you don't need that.”
3
8
1,242
We’re heading to COLM 2026 in San Francisco! Turing will be at Booth 305, bringing together ideas and conversations around the rapidly evolving world of language models. @COLM_conf is a unique academic forum dedicated to understanding, improving, and critically examining language modeling from foundational research to the broader questions shaping the future of AI. If you’re attending COLM 2026, come meet the @Turingcom team. We’d love to connect with researchers, practitioners, and fellow AI enthusiasts exploring what’s next in language modeling. Where: San Francisco, Booth 305 See you there!
4
10
753
Turing retweeted
We're building a strong research team at @turingcom. Hiring at all levels to work on: - The science of data generation and RL environments - Post-training and RL - New benchmarks and evals for frontier models You'll publish and collaborate with academia. DM if interested.
17
24
484
44,766
Turing retweeted
The right name for what the frontier labs are building. Proud to help them build it. @turingcom
2
7
45
2,039
We’re hiring at @Turingcom!
We're building a strong research team at @turingcom. Hiring at all levels to work on: - The science of data generation and RL environments - Post-training and RL - New benchmarks and evals for frontier models You'll publish and collaborate with academia. DM if interested.
3
7
92
12,274
New research from Turing, in collaboration with Oracle: PLSQLBench, a benchmark for evaluating whether LLMs can write executable PL/SQL programs. Most benchmarks test general code generation or text-to-SQL. But real database work isn't a one-shot query, it's stored procedures, cursors, exception handling, and iterating against a live schema. That entire layer has gone basically untested. Until now. Introducing PLSQLBench:  the first benchmark built for procedural database programming. -2,865 instances -2,594 single-turn tasks + 271 multi-turn conversations (978 turns) Built on enterprise-style Spider 2 databases, schema-grounded Spider tasks, and MBPP-derived procedural problems. It measures what developers actually do: write procedures, handle exceptions, ground work in real schemas, and refine iteratively. If we want LLMs that ship real database code, we have to evaluate them on real database work. Testing eight LLMs surfaced recurring weaknesses in schema grounding, PL/SQL dialect fidelity, procedural control flow, exception handling, and consistency across turns. Tool-augmented agents closed some of the gap on schema-grounded tasks, but meaningful gaps remain. The paper has been accepted to EMNLP industry track 2026 in Budapest. Paper + Code below. Congratulations to all: Marianne Menglin Liu, Leo (Leonid) Boytsov, Daniel Petersen, Pramuditha Perera, Rongguang Wang, Sai Ashish Somayajula, Syed Hamza Rafique, Rohit Saini, Shubham Pathak, Sujeeth Bharadwaj, Tao Sheng, Graham Horwood, Fahad Shah, Ankan Bansal, Sujith Ravi, Dan Roth
15
7
32
6,167
Turing retweeted
How do deeptech startups build for longevity in a rapidly evolving AI landscape? Nasscom’s 18-startup #InnoTrek2026 delegation visited StartX - Stanford University's non-profit, zero-equity founder community (home to 1,500+ startups and 30+ unicorns) - to unpack what it takes to scale global ventures. 3 Strategic Takeaways: • Build for tomorrow's AI, not today’s gaps: Avoid merely "plugging current holes" in AI models. Focus on core value propositions that compound as foundation models advance. • Embrace AI as an empirical science: Traditional playbooks are shifting. Success demands hands-on experimentation with frontier models and rapid adaptation to emergent properties. • Leverage cross-border networks: Bridging Indian tech talent with Silicon Valley’s mentors, strategic partners, and investors accelerates global market reach. A powerful exchange reinforcing India’s role in shaping the global deeptech ecosystem #DeepTech #IndiaUS #Startups #Innovation #InnoTrekUSA2026 @StartX @turingcom @doshikavita @nasscomdeeptech
2
9
680
Turing retweeted
SciCode++ is here, with ~6,000 tasks that test how well frontier models solve scientific problems through code. SciCode introduced executable scientific coding problems. SciCode-Verified showed how often flawed specifications and tests can distort the results. SciCode++ turns those lessons into a production process: domain experts write the tasks, independent reviewers check them, executable tests validate every subproblem, and model runs calibrate difficulty. That helps separate model capability gaps from eval issues. Early results post-training Qwen 3.5 9B on 4,000 of the 6,000 tasks: +10.3% on SciCode and +9.1% on SciCode-Verified relative to baseline. turing.com/blog/scicode-plus…
3
11
22
1,079
Turing retweeted
Turing welcomes Patrick McKinney as Chief Information Security Officer. As AI becomes more capable and more deeply embedded in how companies operate, security, privacy, and responsible data stewardship have never been more important. As we scale our work with frontier AI labs and Fortune 500 enterprises, Patrick will lead our global security strategy with a deeply technical, hands-on approach. Patrick will help set the security bar for our Frontier AI work with leading labs on training data and evals and drive new security research and benchmarks like CyberStrike. He’ll also work directly with enterprise customers and their security teams as we deploy agents in highly regulated environments, helping answer the hard security questions that come with putting AI into production. Across this work, he’ll help protect the data entrusted to us and strengthen our security capabilities across cloud identity, detection engineering, incident response, and offensive and defensive security. Patrick brings more than 15 years of experience building and scaling security organizations, including leadership roles at Invisible Technologies and security and compliance roles at Dropbox and Coinbase. Welcome to Turing, Patrick!
2
7
14
1,188
Turing retweeted
Last week, the Turing Frontier Research Lab launched KernelQuest, a benchmark that tests whether AI agents can perform the full job of GPU kernel optimization. KernelQuest contains 100 engineer-authored Triton tasks across 21 kernel families. Each task gives an agent a live environment where it can profile a PyTorch workload, write and revise kernels, and measure performance. The tasks range from single operators and fused workloads to complete models and large Transformer/MoE workloads. Learn more in the article below:
11
7
26
976
Turing retweeted
Proud to see our CTO, @ecekamar named one of the Top Women in AI 100 for 2026. Ece’s leadership continues to shape what responsible, human-centered AI can become. We’re proud to see her recognized among the women building the future of AI. Congrats, Ece!
3
6
22
3,332
Turing retweeted
What drives AI performance beyond the leaderboard? We’re joining Foothill Ventures, EchoHer, Chargebee, and Pillsbury Winthrop Shaw Pittman LLP for an intimate, closed-door conversation with approximately 50 AI founders, builders, and researchers. The discussion will explore what it takes to build AI systems that perform and businesses that scale, including: - Models vs. systems - Production evaluations - Reliability - Inference economics - Durable product advantage @Turingcom’s own Charlotte Tao, Principal, Frontier AI Solutions, will join: -Vinay S., Senior Director of Product at Chargebee -Lei Zhang, Founder and CEO of Stardust AI The conversation will be moderated by Theresa Dai of Foothill Ventures. Curated guests. Focused topics. Thoughtful conversations with founders, builders and researchers. RSVP below.
2
8
13
858
Our Turing Frontier Research Lab recently released CyberStrike, a real-world cybersecurity benchmark. A long-horizon benchmark evaluating frontier agents across the cybersecurity lifecycle, from vulnerability exploitation and secure remediation to threat detection and incident reconstruction. Results across 3,600 trials: GPT-5.6 Sol (xhigh) led at a 31.5% mean per-task pass rate, but many high-scoring near misses failed binary grading. Seventy-two of 200 tasks were unsolved. CyberStrike has 200 expert-authored tasks: 120 defensive, 68 offensive, and 12 DFIR. Of 3,600 recorded trials, 3,106 were graded; the remaining 494 were reported safety refusals. On the atlas-records-export-offensive task, GPT-5.6 Sol reported: "The submitted proof replayed successfully: all 14 HTTP steps returned 200." The verifier found no protected artifact. No configuration solved the task; four declared success, and two refused. At the Turing Frontier Research Lab, we continue building benchmarks, training sets, and training recipes that push frontier models forward in the domains where real capability matters. More on CyberStrike: labs.turing.com/benchmarks/c…
3
3
25
3,198