Journalism • Marketing • AI • @thehypedotnews

San Francisco, CA
Wow
new flux.3 costs 2.8x more than seedream 5.0 flash. is it worth it? the setup: one prompt per task. one shot, no reruns, no edits, same aspect ratio on both sides models: @bfl_ai flux.3, @BytePlusGlobal seedream 5.0 flash tasks: mantis macro, y2k tokyo fisheye snap, fluffy plush toy, chrome art-toy head, thermal portrait price per image #1 seedream 5.0 flash – $0.018 #2 flux.3 – $0.05 the takeaway: flux.3 costs 2.8x more for a stronger camera look. for high-volume work, seedream wins on price follow @thehypedotnews for 24/7 ai news, analysis and breakdowns
1
15
Need for speed
gpt-6.1 sol vs gpt-6 sol vs gpt-5.6 sol – three need for speed style tracks each, built file by file the setup: one prompt per track. the model plans first – track layout, race choreography, shot list, a file manifest with no file over ~350 lines – then writes the project one file per request, with everything it already wrote in context. our own harness on @OpenRouter, 100k tokens max per reply, reasoning effort high for all three, a file that doesn't fit gets continued from its last line every track is a folder of plain es modules, @threejs 0.170 from a cdn, every mesh and texture generated in code – no models, no images, no hdrs. four sports cars racing on the road, four camera shots of 5 seconds each with hard cuts, number keys to jump, space to pause tasks: 1. alpine pass – a hairpin road up the dolomites in golden autumn, an avalanche gallery with arched openings, an old stone bridge over a gorge with a waterfall, larches, cows, a chapel and a cable car 2. coast road – a cliffside road above a turquoise sea, a stone viaduct over a cove, rock tunnels, a terraced pastel village with a tiled dome, a harbor and a lighthouse 3. neon city – a rain-soaked night race on an elevated expressway and through a street canyon of neon signs, wet asphalt mirroring the lights, a harbor and a lit bridge every track has the same four shots: an aerial overview, a roadside pass-by, a cockpit view from the lead car and one free shot unique to the track all 9 tracks render. every track cost under $2, and gpt-6 sol is the cheapest on all three total cost, three tracks #1 gpt-6 sol – $3.90 #2 gpt-6.1 sol – $5.00 #3 gpt-5.6 sol – $5.63 cost per track, 6.1 sol / 6 sol / 5.6 sol alpine pass – $1.44 / $1.39 / $1.87 coast road – $1.91 / $1.16 / $1.93 neon city – $1.65 / $1.35 / $1.82 time, three tracks #1 gpt-5.6 sol – 1h 11m 11s #2 gpt-6 sol – 1h 18m 37s #3 gpt-6.1 sol – 1h 39m 2s total output tokens #1 gpt-6 sol – 286,519 #2 gpt-6.1 sol – 358,667 #3 gpt-5.6 sol – 360,684 lines of code shipped #1 gpt-6 sol – 18,605 #2 gpt-6.1 sol – 22,009 #3 gpt-5.6 sol – 25,143 observations: • gpt-6.1 sol follows the brief closest. its alpine pass stacks seven hairpins under the dolomite towers, with a waterfall and a rainbow in the spray next to the arched bridge. its neon city has the wettest asphalt of the grid, with every sign and headlight smeared across the road • gpt-6 sol is the cheapest on every track and the leanest thinker – 287k output tokens for all three. its coast road puts a terraced village with red roofs above a turquoise cove, and its neon expressway runs through rain and paper lanterns • gpt-5.6 sol ships the most code, 25k lines, in the shortest request time. its cars are the most detailed inside: shift paddles, round vents, a lit instrument cluster and a live rear-view mirror follow @thehypedotnews for 24/7 ai news, analysis and breakdowns
23
new bench
new sonnet 5.5 leaves sonnet 5 far behind – three one-shot silent hill scenes the setup: one prompt, one reply, no agent loop. every scene is one html file with three.js from a cdn. every model is built from code, with no model files and no images. tasks: 1.pyramid head 2.the church 3.the nurse models: @AnthropicAI's sonnet 5.5, sonnet 5 all 6 scenes render. sonnet 5.5 spends more on every scene and puts the extra into the frame: generated textures, soft shadows, fog, ash particles and a full otherworld transition on two of the three scenes tokens: #1 sonnet 5 – 2,794 #2 sonnet 5.5 – 20,141 time: #1 sonnet 5 – 36s #2 sonnet 5.5 – 6m 14s lines of code: #1 sonnet 5.5 – 1,260 #2 sonnet 5 – 659 observations: - sonnet 5.5 builds a scene, not only a model. pyramid head gets a procedurally textured helmet with rivets, a knife that leaves a drag mark on the floor, aces tone mapping and soft shadows, all generated in code - its church is the only one of the two with a graveyard, trees and a fence around it. flip the otherworld switch and the fog, sky, lights and stained glass fade into rust-red instead of cutting - the nurse is the most alive of the six: a slow sway, breathing and random head jerks, all procedural, no animation files - sonnet 5 is the fast sketch: about 12 seconds and under 1k tokens per scene, clean readable geometry and a working turntable every time. all three scenes took it fewer tokens than sonnet 5.5 spent on any single one - the whole sonnet 5.5 grid fits in about 20k tokens and just over 6 minutes. that is 7x the tokens for a visibly different class of build follow @thehypedotnews for 24/7 ai news, analysis and breakdowns
1
28
this is extra cozy
claude sonnet 5.5 vs gpt-6 sol the task: build a 3d apartment that runs in a web browser – a kitchen, a living room, a bedroom and a real view out of the window. no ready-made 3d models, no photos: every chair, every texture, every lamp is written in code. the apartment plays itself like a real-estate video, the camera flies through each room. three apartments per model the setup: same brief for both models, run through our own harness on @OpenRouter. the model writes a plan first, then the code file by file. headless chrome is the only judge – it loads the build, flies through all six camera shots and hands back a frame from every one. @threejs 0.170 from a cdn. reasoning medium for sonnet 5.5, high for sol apartments: • industrial loft – 58th floor in manhattan at golden hour. a brick wall, copper pots on a brass rail, an espresso machine with a steaming cappuccino, a record corner with glowing amp tubes, a cat tree, and the city lighting up behind a black steel window grid • eco cabin – the alaskan forest at blue hour. a river-stone wall with a live wood stove, a pour-over coffee station, sourdough on a floured board, and an aurora over snowy spruces through a triangular window • minimalist penthouse – miami beach, late morning. a travertine island, an infinity pool on the terrace, sheer curtains in the sea breeze, a surfboard by the glass, sailboats and pelicans over the atlantic models: @AnthropicAI claude sonnet 5.5 @OpenAI gpt-6 sol all 6 apartments render. sonnet 5.5 ships more code on two of the three and plans the finest. sol is the cheaper and the faster on every apartment total cost, three apartments #1 gpt-6 sol – $5.13 #2 claude sonnet 5.5 – $21.80 cost per apartment, sonnet 5.5 against sol loft – $8.80 vs $1.89 eco – $7.37 vs $1.75 minimal – $5.64 vs $1.50 generation time, three apartments #1 gpt-6 sol – 1h 18m 55s #2 claude sonnet 5.5 – 2h 44m 03s total tokens #1 gpt-6 sol – 351,710 #2 claude sonnet 5.5 – 1,398,777 lines of code shipped #1 claude sonnet 5.5 – 26,317 #2 gpt-6 sol – 20,863 observations: 1. claude sonnet 5.5 plans like an architect: a placement table with position, size and support surface for every object. it split the eco cabin and the penthouse into 75 and 76 modules, and its eco cabin is the biggest build of the six at 10,709 lines 2. sonnet 5.5 is the one to watch for materials: green hand-glazed tiles, fire flickering through the stove glass, a wall of river stones, a travertine island with its veining running over the edge 3. gpt-6 sol built each apartment in under 30 minutes for under $1.90, in about 30 files each. its loft has a library ladder, a bike on the brick wall and a desk looking out over manhattan. its eco cabin ends on a cedar hot tub steaming under the aurora 4. sol costs 3.8x to 4.7x less than sonnet 5.5 per apartment and is 1.6x to 2.7x faster on each one six apartments, three styles, about $27 in total follow @thehypedotnews for 24/7 ai news, analysis and breakdowns
2
33
robots
5 open-source robots you can build yourself 3d print files, parts lists and code are free. sorted from easiest to hardest 1. reachy mini, from $299 a desktop robot by @pollenrobotics and @huggingface. it looks around, listens and shows emotions with its antennas. a kit, 2–3 hours to build github.com/pollen-robotics/r… 2. so-101, from $122 a robot arm by @therobotstudio and @huggingface. move one arm by hand and the other copies you, then an ai learns the task on its own github.com/therobotstudio/so… 3. xgo duck, about $410 in parts a walking robot duck by @luwu_dynamics on an arduino uno q. it walks, gets up after a fall and picks up small objects github.com/luwudynamics/xgod… 4. open duck mini, under $400 a homemade version of disney's bdx droid by @antoinepirrone and the community. it learns to walk in a simulator, just like the original github.com/apirrone/open_duc… 5. lerobot humanoid, about $2,500 in parts 3d-printed walking legs by @huggingface. 12 motors run by a raspberry pi 5, and it learns to walk in a simulator github.com/huggingface/lerob… follow @thehypedotnews for 24/7 ai news, analysis and breakdowns
3
143
rich vibes
glm 5.3 prime vs mimo v2.6 pro vs deepseek v4.1 flash – a yacht, a jet and a race car each, built file by file the setup: one prompt per scene. the model plans first – concept, shot list, a file manifest with no file over ~350 lines – then writes the project one file per request, with everything it already wrote in context. our own harness on @OpenRouter, 32k tokens max per reply, a file that doesn't fit gets continued from its last line. tasks: 1. yacht – a small yacht on animated open water, heave, pitch and roll read from the wave surface under the hull, a furnished saloon and cabins, sky, sun and horizon 2. business jet – a private jet in flight, passes through and past clouds, the ground below, a fully furnished cabin with shots inside it 3. formula car – an open-wheel car lapping a full circuit on a racing line, spinning wheels, steering fronts, a body that reacts to braking and cornering, a detailed cockpit with onboard shots models: @Zai_org glm 5.3 prime, @XiaomiMiMo mimo v2.6 pro, @deepseek_ai deepseek v4.1 flash all 9 scenes render. two of the three models built all three scenes for under a dollar, and deepseek v4.1 flash is the fastest on every task, never by less than 2.9x total cost, three scenes #1 deepseek v4.1 flash – $0.662 #2 mimo v2.6 pro – $0.723 #3 glm 5.3 prime – $15.05 cost per scene, deepseek / mimo / glm yacht – $0.203 / $0.217 / $4.96 jet – $0.248 / $0.194 / $5.10 formula car – $0.212 / $0.313 / $4.99 time, three scenes #1 deepseek v4.1 flash – 1h 8m 54s #2 glm 5.3 prime – 3h 56m #3 mimo v2.6 pro – 5h 50m 31s total output tokens #1 mimo v2.6 pro – 673,483 #2 deepseek v4.1 flash – 993,667 #3 glm 5.3 prime – 1,393,959 lines of code shipped #1 mimo v2.6 pro – 16,407 #2 glm 5.3 prime – 18,370 #3 deepseek v4.1 flash – 30,169 observations: • deepseek v4.1 flash wrote 30k lines across 79 files for 66 cents – each scene in 21 to 25 minutes. its jet alone is 13k lines, with its own flight dynamics and contrail modules; its race car keeps the physics and the visuals in separate modules, with a dedicated racing line • mimo v2.6 pro is the leanest thinker of the three – 673k output tokens for the whole grid – and the cheapest on the jet, $0.194. its yacht sits in a golden-hour sea with a guest cabin, twin berths and a bedside lamp • glm 5.3 prime builds the most furnished interiors: a wood-trimmed yacht with a saloon, a galley and a helm, a jet cabin with rows of club seats, a circuit with grandstands, kerbs and a t-cam over the airbox. it also thinks the longest – with reasoning capped at 20k it wrote its last 37 files in under 10 minutes for $2.81 • a year ago a multi-file three.js project with a furnished interior was a frontier-model job. here two open chinese models ship three of them for less than a dollar each follow @thehypedotnews for 24/7 ai news, analysis and breakdowns
1
41
😎
opus 5.5 vs fable 5.1 vs gpt-6 astra – three post-apocalyptic films, one html file each the setup: the same ~100k-token prompt for every model – a three.js skill pack, the brief and a mandatory delivery checklist. gpt-6 astra got one prompt and one reply on @OpenRouter, no agent loop. opus 5.5 and fable 5.1 ran as agents in claude code: write the file, open it in a browser, look at the frames, fix, repeat. opus tokens come from the session log and are priced at openrouter rates ($4 / $20 per million, cache reads $0.20). fable cost and tokens are estimates from file size, its wall clock is measured every film is one html file, three.js 0.170 from a cdn, every texture generated in code. four shots on an auto-playing timeline, ~90 seconds, letterbox and fades, keys 1-4 jump between shots tasks: 1. amusement park – a fogbound park at sunset: through the gate, a dolly around the carousel, a coaster pov over the first drop, a crane up the ferris wheel 2. ghost town – a mojave town at sunset: highway drone shot, main street dolly, a diner interior lit like a film set, a crane reveal over the whole grid 3. nuclear plant – a night sky as the hero: 30,000+ stars and a structured milky way over two 150 m cooling towers, the workers' town, the control room models: @AnthropicAI claude opus 5.5, claude fable 5.1, @OpenAI gpt-6 astra all 36 shots render. opus and fable land all 12 on brief, astra lands 11 – its plant has almost no stars, and the sky was the whole point of that task total cost, three films #1 gpt-6 astra – $8.43 #2 claude fable 5.1 – ~$9.20 #3 claude opus 5.5 – $19.38 cost per film, opus vs astra amusement park – $4.85 vs $2.54 ghost town – $9.94 vs $3.08 nuclear plant – $4.59 vs $2.81 wall clock, three films #1 gpt-6 astra – 27m 16s #2 claude fable 5.1 – 40m 3s #3 claude opus 5.5 – 69m 18s output tokens #1 gpt-6 astra – 93,905 #2 claude fable 5.1 – ~98k #3 claude opus 5.5 – 373,076 lines of code shipped #1 gpt-6 astra – 4,714 #2 claude opus 5.5 – 2,580 #3 claude fable 5.1 – 1,691 observations: - opus 5.5 found its own bugs by looking at its own frames: a nan in the light shafts that drew black dashed lines in the diner, "patch" used as a variable name (a reserved word in glsl), a camera that ended the control room shot staring into the console, trees that hid the blinking stack light at the end of the avenue - the loop gets cheaper as it goes. opus built the ghost town first – 46 model calls and $9.94 – then the amusement park in 16 calls for $4.85 and the nuclear plant in 17 calls for $4.59, reusing its own kit from the first film - $8.8 of the opus $19.38 is cache reads – every agent step re-reads a 200k-450k context - gpt-6 astra is the fastest and the cheapest, 7m 18s to 10m 32s per film in one reply with no browser. its amusement park and ghost town run all four shots - only opus put shafts of sun through dusty blinds in the diner and reflected its 38,000 stars in the cooling channel. fable matched it on the sky – 42,000 stars and a milky way with dust lanes follow @thehypedotnews for 24/7 ai news, analysis and breakdowns
2
37
new bench. star wars 😎
gpt-6 sol vs grok 4.7 vs gpt-6 astra vs muse spark 1.3 – three star wars worlds each, one html file per world the setup: one prompt, one reply, no agent loop. our own harness on @OpenRouter, headless chrome as the only judge – it loads the file, presses 1 / 2 / 3 / 4 and hands back a screenshot of every shot plus every console error. no rubric, no note from us. a file that crashed got one more turn with the console pasted back. grok 4.7 got extra turns on top, by our request – art direction as numbers, never a line edited by hand every world is one html file, three.js 0.170 from a cdn, every texture generated in code – no model files, no images. an auto-playing cinematic with a shot timeline, letterbox and fades. reasoning high for astra, muse and sol; grok 4.7 ran at reasoning low tasks: 1. death star – a slow orbital approach with star destroyers for scale, the corridor into the throne room, the station over a planet's horizon, then the superlaser: eight tributary beams, one green shot, a shockwave ring and thousands of instanced fragments 2. coruscant – the planet from orbit as a circuit board of glowing hubs, a daytime flythrough between towers, under a bridge and past the senate dome, a neon night district with a highway of light trails 3. kamino – an ocean world under storm clouds, tipoca city on stilts with gerstner waves, gpu rain and lightning, a glass walkway over a hall of marching clones two hard rules in every brief: zero console errors on the first run, and nothing loaded from outside the file except three.js itself. models: @OpenAI gpt-6 sol, @xai grok 4.7, @OpenAI gpt-6 astra, @AIatMeta muse spark 1.3 all 12 worlds render. sol is the fastest on every one of the three tasks and never by less than 2x, and the whole grid came in at $6.99 total cost, three worlds #1 muse spark 1.3 – $0.385 #2 gpt-6 sol – $0.556 #3 grok 4.7 – $1.732 #4 gpt-6 astra – $4.314 cost per world, sol against astra death star – $0.202 vs $1.476 coruscant – $0.190 vs $1.559 kamino – $0.164 vs $1.279 wall clock, three worlds #1 gpt-6 sol – 6m 23s #2 muse spark 1.3 – 14m 27s #3 gpt-6 astra – 26m 52s #4 grok 4.7 – 58m 20s total tokens #1 gpt-6 sol – 59,299 #2 gpt-6 astra – 89,967 #3 muse spark 1.3 – 102,716 #4 grok 4.7 – 499,794 lines of code shipped #1 gpt-6 astra – 4,274 #2 grok 4.7 – 4,034 #3 gpt-6 sol – 3,046 #4 muse spark 1.3 – 1,664 observations: - gpt-6 sol shipped all three worlds on the first reply, 2m 21s or less each, zero console errors, 4k to 6k reasoning tokens per file. its coruscant is the only daytime city of the four with a bridge between the towers and traffic at three altitudes - grok 4.7 is the model that takes direction. we gave it the exact camera path for tipoca city – 19 points with 2.5 units of clearance from every dome – and it landed the flythrough in one edit. we pointed at a one-character bug in its night window texture and it lit the whole district in one edit. 12 rounds across three worlds and not one console error in any of them - gpt-6 astra is the closest to the film frames on the death star: a grey station with a dark side, the ring-walled throne room, a superlaser that cracks the planet before it blows. it is also 7.8x sol on cost - muse spark 1.3 is the cheapest on every task and never by less than 1.2x against sol – $0.385 for the grid, 11x under astra - sol did three worlds in 6m 23s, less than astra spent on any single one follow @thehypedotnews for 24/7 ai news, analysis and breakdowns
3
48
grok 4.7
grok 4.7 vs gpt-6 astra in 3d scenes we put the two models on one job: three cinematic three.js scenes, each a single self-contained html file, no textures, no models, no libraries beyond three.js. a rocket launch with stage separation, a meteor strike that knocks down a city, a dam break that washes away a town. gpt-6 astra – @OpenAI, $5/$25 per m tokens on the flex tier via @OpenRouter grok 4.7 – @SpaceXAI, shipped sep 21, $1.6/$4.8 per m tokens via @OpenRouter - cost, three scenes #1 grok 4.7 – $1.31 #2 gpt-6 astra – $4.11 - tries #1 gpt-6 astra – 12 #2 grok 4.7 – 15 - wall clock, all calls summed #1 gpt-6 astra – 25m 14s #2 grok 4.7 – 60m 43s - output tokens spent on thinking #1 gpt-6 astra – ~34% (14–17k per scene) #2 grok 4.7 – ~84% (62–86k per scene) observations: • grok thinks for 17–23 minutes per scene. 79k of its 93k rocket tokens were reasoning. a streaming request goes silent that long: one attempt was cut off and two more hung with zero bytes before a plain non-streaming call got through • astra's first rocket had a concrete plain, we asked for grass, and a square bloom halo around the distant stage – both fixed by asking, not by editing. its meteor only switched the city lights off; "the buildings must physically collapse" was one more request • grok's dam break rendered as a white ball. 6,200 spray particles drawn additively on top of a bloom pass, fixed by hand grok 4.7 is 3x cheaper per scene and 2.4x slower
2
51
benchmarked union alpha
union alpha (unbiased pareto) vs deepseek v4.1 flash vs muse spark 1.3 – three paintings in three.js the setup: one four-line prompt plus the painting as an image, through @openrouter. no agent loop, no renders, no feedback – the model writes one html file blind and we open it. @threejs from a cdn, every texture generated in code. when a provider cut the stream early we sent the partial back and said continue exactly where you stopped tasks: 1. the starry night – van gogh, 1889 2. the persistence of memory – dalí, 1931 3. poppies at argenteuil – monet, 1873 two rules in every brief: keep the painting's palette, brushwork and mood, and reply with the code only models: @theunbiasedco union alpha (stealth, free), @deepseek_ai deepseek v4.1 flash, @aiatmeta muse spark 1.3 total cost, three scenes #1 union alpha – free (list price: $1.04) #2 muse spark 1.3 – $0.117 #3 deepseek v4.1 flash – $0.162 generation time, three scenes #1 muse spark 1.3 – 4m 45s #2 deepseek v4.1 flash – 11m 30s #3 union alpha – 29m 13s total completion tokens #1 muse spark 1.3 – 26,502 #2 union alpha – 123,756 #3 deepseek v4.1 flash – 156,028 lines of code shipped #1 muse spark 1.3 – 862 #2 union alpha – 1,715 #3 deepseek v4.1 flash – 2,816 observations: • union alpha reads the painting like an art historian. it named every work unprompted, then broke each into parts: dalí's watches deformed along a bezier curve, monet's poppies as instanced brush dabs under a wind shader. no other model went that deep on a one-shot • it is the only model that made the paintings move the way they were painted. in the monet, the woman and the child walk the field on catmull-rom paths, pollen drifts, poppies are brush dabs in a point shader that sway in the wind. deepseek and muse left the figures standing • its dalí is an inventory of the canvas: three soft clocks draped along one parametric curve, a drip falling off the hanging one, a fly, ants on the pocket watch as an instanced mesh. 17 named parts in all. nobody else drew the drip • its starry night is shader work end to end: shared glsl noise, a vortex field for the sky, billboarded shader quads for the moon and stars, a painterly surface shader for the hills, and windows that flicker on their own timers. the file reads like a demoscene entry, not a model output what union alpha is: • we asked it. the stealth window had closed a day after launch and the api answered: "this model was unbiased's pareto". pareto is from circuit & chisel, an ex-stripe team that raised $19.2m in sep 2025 per fortune, and now sells "frontier intelligence for 75% less" • pareto is not one model. per unbiased's site it "runs a mix of frontier and open source models against each other on every request" and keeps the best answer. that is the 300-second first token, the tokenizer listed as "other", and the missing reasoning field – a race, not a model • listed at $2.50 in and $7.50 out per million. our three scenes would have cost $1.04 – 6.4x deepseek, 8.9x muse. asked for its cutoff, it dated nothing past may 2025 follow @thehypedotnews for 24/7 ai news, analysis and breakdowns
1
98
DeepSeek 🐳🐳🐳
deepseek v4.1 flash ran all four three.js scenes for $0.15. gemini 3.8 flash needed $1.25 we put the two models on one job: a self-contained html page that renders an animated 3d scene with three.js – a storm at sea, a spiral galaxy, a tornado over farmland, a tesla coil on a dark table. same prompt - calls until all four scenes were usable #1 deepseek v4.1 flash – 6 #2 gemini 3.8 flash – 12 - time #1 deepseek v4.1 flash – 13m 57s #2 gemini 3.8 flash – 32m 02s - cost #1 deepseek v4.1 flash – $0.15 #2 gemini 3.8 flash – $1.25 observations: • gemini blew out the exposure on 4 of 5 first attempts – a white blob where the tesla coil should be. the prompt had a paragraph on exposure. it did not help. • both models overcorrect in the fix loop: "too bright" becomes "almost black". round two is where it lands. same four scenes, 6 calls and $0.15 against 12 calls and $1.25!
2
61
New one
gpt astra vs fable 5.1 at goldberg machine gpt 6 astra – openai, landed on @OpenRouter less then hour ago, provider pinned to openai fable 5.1 – anthropic, shipped sep 1 we put the two models on one job: a rube goldberg machine in three.js that presses a button and detonates a bomb the setup: one self-contained html file, three.js from a cdn, everything else procedural – no textures, no models, no physics engine, every collision hand-written. the hard part sits in the brief: a domino may only fall once the previous one actually touches it, checked by real overlap every frame, never by a timer. same rule for the hammer hitting the button and the button firing the bomb. one continuous camera, its speed driven by whatever is moving. we recorded both scenes frame by frame – 1200 frames, 60 fps, exactly 20 seconds – and stepped both by hand to read the telemetry. - cost #1 astra – $1.84 #2 fable – $29.16 - time #1 astra – 9m 56s #2 fable – 1h 12m - tokens #1 astra – 45k #2 fable – 360k - lines of code astra – 881 fable – 744 observations: • we told it what we saw and nothing else – no diagnosis, no patch. we never edit a model's code. round two ran the whole chain to the blast. • both files are deterministic. two runs each, identical state to twelve decimals, and neither model reached for math.random. conclusion: 15.8x cheaper and 7.2x faster, and it still took a second round to get the ball into the bucket! follow @thehypedotnews for 24/7 ai news, analysis and breakdowns
2
35
nickster retweeted
GPT Astra just cooked Claude Fable 5.1 — for a lower price! OpenAI just rolled out GPT‑6 Astra — we put it against Anthropic's Fable 5.1 Both models got the same four briefs: Earth colliding with Mars, a tall ship sucked into a giant whirlpool, Godzilla leveling downtown, a mega-tsunami swallowing a city. One self-contained HTML file per scene, one shot, no edits. GPT Astra: 25,159 tokens, $1.67 Claude Fable 5.1: 37,767 tokens, $2.51 both models live on aimlapi[.]com — 1000+ models, one key.
GPT-6 Astra is now available to all Pro, Enterprise, and Business Premium users in ChatGPT Work and Codex. It's also live in the API. It might take a few days to roll out to our Plus and Business users. Thank you for your patience.
15
15
249
90,321
animals benchmark
gemini 3.8 flash vs muse spark 1.3 – three animated 3d animals each, in an agent loop the setup: our own agent loop on @openrouter, a browser as the tool set – write the file, patch it, render it, sign off. the harness loads the scene in headless chrome, presses 1 / 2 / 3 and hands back a screenshot of every shot plus every console error. no rubric, no judge, no note from us 16 turns and 7 renders, hard ceiling. the model also gets two three.js reference documents in its system prompt and can pull deeper reference files on demand. every scene is one html file, three.js from a cdn, every texture generated in code – no model files, no images tasks: 1. lizard – sprinting a jungle trail. diagonal-couplet gait, a lateral wave down the spine, forked tongue on the close shot 2. macaw – over the canopy. a real flap cycle: primaries closed on the downstroke, wrist folded and split on the upstroke 3. tiger – crouch, leap to a branch, walk it, curl up and sleep at dusk with moths glowing around it. the leap has to be ballistic two hard rules in every brief: the animal fills at least 40% of frame height in every shot, and the body is one continuous surface swept along the spine, not a stack of capsules. models: @googledeepmind gemini 3.8 flash, @aiatmeta muse spark 1.3 all six scenes render. muse is cheaper on every one of the three tasks and never by less than 1.6x, and the whole grid came in at $4.32 total cost, three scenes #1 muse spark 1.3 – $1.488 #2 gemini 3.8 flash – $2.836 cost per scene, muse against gemini lizard – $0.551 vs $0.897 macaw – $0.456 vs $0.877 tiger – $0.481 vs $1.062 wall clock, three scenes #1 muse spark 1.3 – 40m 46s #2 gemini 3.8 flash – 58m 03s total tokens #1 muse spark 1.3 – 1,461,756 #2 gemini 3.8 flash – 1,794,173 lines of code shipped #1 muse spark 1.3 – 2,216 #2 gemini 3.8 flash – 4,736 observations: • muse asked for more renders while spending half the money – 19 against 17 – so the extra spend on gemini's side is not extra looking, it is extra writing • none of the six runs called finish. all six hit the 16-turn ceiling, so neither model was ever satisfied with what it saw • gemini shipped the tiger broken. it rewrote the scene on turn 14, wrote "ry0 is not defined" into it, saw the exception in its own render on turn 16 and ran out of turns. four more turns and it fixed it in one edit • we ran the same three briefs one-shot first, blind, with no renders and no feedback – $0.38 for gemini and $0.35 for muse. both shipped stacks of capsules butted end to end. the loop costs 5.9x more and it is the only reason any of these reads as an animal • without the 40% rule both models read "wide shot from the branches" literally and put a 4-pixel speck in the middle of a landscape follow @thehypedotnews for 24/7 ai news, analysis and breakdowns
2
43
this is amazing
One home. 4 distinct spaces. A Garden, Living Room, Dining Room and Bedroom morph seamlessly through procedural geometry, animation and dynamic lighting. Inspired by a demo from my talented friends at @threejsassets . I can open source the code if there is enough interest.
1
7
213
fable 5.1 being tolkien
fable 5.1 vs fable 5 vs opus 5 – three lord of the rings landmarks, built in 3d from one image the setup: one reference image per scene, one html file per build, everything procedural – no meshes, no textures, no image files, nothing past @threejs from a cdn. each model reads the picture, writes its own prompt from it, then builds to that prompt in the same turn. three named camera shots per scene on keys 1/2/3, so it can be screen-recorded. run through @openrouter tasks: 1. bag end – hobbiton from two frames, outside and in. the round green door has to open onto the room you are standing in 2. barad-dûr – the tower and orodruin from one film still. the eye has to move and track the camera, the volcano erupts on a cycle, the clouds never stop 3. rivendell – jerry vanderstelt's painting. sun shafts that shimmer, water that falls without a break, trees that sway on a gust models: @anthropicai fable 5.1, fable 5, opus 5 total cost, three builds #1 fable 5 – $14.97 #2 opus 5 – $18.53 #3 fable 5.1 – $22.38 wall clock, three builds #1 fable 5 – 38m #2 fable 5.1 – 92m #3 opus 5 – 122m output tokens #1 fable 5 – 298,592 #2 fable 5.1 – 439,435 #3 opus 5 – 724,418 lines of code shipped #1 fable 5 – 2,885 #2 fable 5.1 – 4,021 #3 opus 5 – 5,161 biggest single build, lines #1 opus 5, bag end – 2,410 #2 fable 5.1, barad-dûr – 1,375 #3 fable 5, bag end – 1,319 observations: • fable 5.1 is the only model that furnished the bag end interior – a live fire, panelling, books on the floor, leaded diamond windows, against fable 5's flat color and opus's dark tunnel. the round door outside opens onto that room, the hard part of the brief • what it costs is thinking room. the 128k output ceiling is a thinking budget in disguise: fable 5.1 burned 102,116 of it on reasoning and hit the wall mid-file. opus spent 109,241 and hit the same wall. fable 5 spent 61,240 and finished bag end in one call – the only one that did • fable 5.1's first pass is not the finished thing. its barad-dûr came back with three defects you only catch by looking at it – nothing a read of the code would have flagged • it is the best of the three at being corrected. handed a plain list of what was wrong, it returned 32 targeted patches over two rounds, every one applied first try, and it worked out one of the causes itself instead of guessing at constants conclusion: nine scenes, 12,067 lines and 1.46m output tokens for $55.88 all in – and the cheapest model was also the fastest, by 3.2x! follow @thehypedotnews for 24/7 ai news, analysis and breakdowns
6
57
claude fable 5.1 is a bee
one prompt, 43 minutes, $2.96 – claude fable 5.1 wrote a 1,660-line world we gave it one job: build a detailed frutiger aero world in a single self-contained html file, with a camera travelling through it for 600 frames the setup: @threejs 0.169 and nothing else – no images, no downloaded textures, no physics library. every surface drawn in code. the file has to expose reset/step/simstate, render 600 frames and produce the same picture twice. one pass. every number below is ours, measured on the rendered frames what came back: • 1,660 lines in one file, 61,363 characters • 60,000 instanced blades of grass, 500 flowers • 400 towers, 700 canopy pieces, 90 rocks • 60 bubbles in the system, 30 on screen at frame 600 • 1.29m triangles at 47 draw calls • a 132m camera path across 8 control points • rendered at 2x and downsampled observations: • it seeded its generator with 20050607 – june 7, 2005, the year the look it was copying belongs to • 47 draw calls for 1.29m triangles. it instanced everything that repeats, without being told which • the brightest frame is 3.6% pure white. the brief capped it at 25%, and this look usually blows out conclusion: one file, no assets, 1.29m triangles and a camera that walks 132 metres through it, for $2.96! follow @thehypedotnews for 24/7 ai news, analysis and breakdown
1
5
69
built some skyscrapers
hy4 preview vs glm 5.3 vs qwen 3.8 max hy4 shipped all three for $0.537 – 3.2x under qwen. glm broke the fewest times, 2 fixes against hy4's 5. qwen wrote the most code, burned the most tokens, and still ships one of its three builds broken the setup: three prompts of ~600 lines each, one html file per build, everything procedural – no meshes, no image files, no libraries past @threejs prompts: 1. the petronas twin towers in dawn mist 2. taipei 101 in a tropical rainstorm 3. the cn tower in heavy snow each file carries the geometry, the weather, night lighting, five scripted camera shots, a capture mode and a self-check panel that prints its own numbers. run through @OpenRouter, one shot per model, no reference images – two of the three models have no vision at all models: @TencentHunyuan hy4 preview, @Alibaba_Qwen qwen 3.8 max, @Zai_org glm 5.3 - total cost, three builds #1 hy4 preview – $0.537 #2 glm 5.3 – $1.111 #3 qwen 3.8 max – $1.737 - wall clock, three builds #1 glm 5.3 – 44m #2 hy4 preview – 50m #3 qwen 3.8 max – 102m - output tokens #1 hy4 preview – 186,635 #2 glm 5.3 – 231,481 #3 qwen 3.8 max – 248,655 - fixes needed to make it run #1 glm 5.3 – 2 #2 qwen 3.8 max – 4 #3 hy4 preview – 5 - lines of code shipped #1 hy4 preview – 1,798 #2 glm 5.3 – 2,766 #3 qwen 3.8 max – 3,066 observations: • every bug was one or two lines. no model failed the architecture – the geometry, the camera rigs and the self-check math were right everywhere. they broke on things a single run catches • glm's petronas attempt spent 148,574 output tokens on reasoning and emitted zero characters of code. capping its thinking budget at 26k re-ran the same prompt in 683s for $0.250 – 3.2x faster and 2.7x cheaper • no model won two towers in a row. taipei went to glm, petronas to qwen, cn tower back to glm, and the failures move the same way. the spread between tasks is bigger than the spread between models conclusion: nine towers, 7,630 lines and 666,771 output tokens for $3.38 all in – and the cheapest model got there on 41% less code than the priciest! follow @thehypedotnews for 24/7 ai news, analysis and breakdowns
4
65
made this one 😎
glm 5.3 vs qwen 3.8 vs gemini 3.7 vs deepseek v4 flash four models designed and built three structures each on a physics-backed site, with no dimensions anywhere in the brief the setup: our own agent loop on @OpenRouter, a construction site as the tool set – footings, walls, arches, roofs, scaffold, a lamp. the site enforces physics and nothing else: unsupported brick falls, a roof needs walls under it, a worker reaches 3.2 m above whatever he stands on, an arch needs centring until the keystone is set, concrete cures before it carries. no budget ceiling – material cost is tallied and reported, never blocked. tasks: 1. house – a plot and a palette, no plan. shape, height and material are the model's call 2. lighthouse – a headland cut by a gully, with a rock stack standing 30 m offshore. the lamp must burn, it must be the highest thing built, and the keeper must be able to walk to it 3. bridge – a river with one islet and banks at different heights. cross it however you want models: @Zai_org glm 5.3 flash, @Alibaba_Qwen qwen 3.8 flash, @GoogleDeepMind gemini 3.7 flash, @deepseek_ai v4 flash vision all twelve objects were finished and signed off by the models themselves. tallest lighthouse is qwen's at 38.4 m, planted on the offshore stack with a bridge run out to it – the only model that read the site that way. deepseek signed off its bridge on an empty riverbed: 0 bricks, 107 minutes, $1.16m of material tallied - total cost, three builds #1 glm 5.3 flash – $0.201 #2 gemini 3.7 flash – $0.871 #3 qwen 3.8 flash – $1.058 #4 deepseek v4 flash – $1.567 - wall clock, three builds #1 gemini 3.7 flash – 91m #2 glm 5.3 flash – 228m #3 deepseek v4 flash – 502m #4 qwen 3.8 flash – 912m - total tokens #1 gemini 3.7 flash – 3,567,052 #2 glm 5.3 flash – 4,732,748 #3 qwen 3.8 flash – 13,469,333 #4 deepseek v4 flash – 18,230,076 - defects logged by the site #1 deepseek v4 flash – 59 #2 gemini 3.7 flash – 132 #3 glm 5.3 flash – 221 #4 qwen 3.8 flash – 350 - material tallied across three builds #1 gemini 3.7 flash – $359,884 #2 glm 5.3 flash – $583,358 #3 deepseek v4 flash – $1,327,484 #4 qwen 3.8 flash – $2,188,625 observations: • glm is the cheap one and nothing here is close – $0.201 for three buildings, $0.042 per million tokens, 6x under gemini's rate • what glm spends it on is bulk, not care: 166,228 bricks in one house and 156 defect weight, the worst single object in the set • gemini is the efficiency line – 91 minutes and 3.57m tokens for all three and an eighth of qwen's clock • gemini also builds the smallest of everything. its lighthouse is 22.5 m against qwen's 38.4, its house 6.9 m against 19.3 • qwen is the maximalist: 1.18m bricks, $2.19m of material, tallest on all three tasks, and 912 minutes – 15 hours – to get there conclusion: twelve finished objects for $3.80 all in, and a 7.8x price spread between the cheapest model and the priciest! follow @thehypedotnews for 24/7 ai news, analysis and breakdowns
1
37
glm 5.3 flash ⚡️⚡️⚡️
glm 5.3 flash is 7.5x cheaper, but 3.4x slower than gemini 3.7 flash @Zai_org glm 5.3 flash – shipped aug 26, $0.07/$0.25 per 1m @GoogleDeepMind gemini 3.7 flash – shipped aug 13, $0.38/$1.88 per 1m we put the two models on one job: write one html file that draws an animated 3d scene in the browser. no images, no downloads, and it has to look the same on every load. the setup: three scenes – a glass aquarium in a lit room, the solar system, a night city under a thunderstorm. identical brief word for word, reasoning effort high, 64k output cap. the numbers below are not the whole run. they cover the three scenes we kept – the best one per task from each model, the ones in the video. - total generation time for the three scenes #1 gemini 3.7 flash – 10m 36s #2 glm 5.3 flash – 36m 30s - tokens spent on those three scenes #1 glm 5.3 flash – 110k #2 gemini 3.7 flash – 111k - cost of those three scenes #1 glm 5.3 flash – $0.027 #2 gemini 3.7 flash – $0.202 observations: • glm's first 10 attempts: 7 blank pages. it kept inventing short random helpers and forgetting to define one of them. the fix was one line in the brief: use exactly one random helper, named rand(), and don't invent shorthands next to it. next 12 attempts: 11 alive, 0 crashes. • glm spends 66% of its output on reasoning, gemini 57%. that is the whole speed gap. • gemini's storm came back as a black rectangle in 4 of 6 runs. glm's best storm has a branching bolt, lit rain and wet asphalt – for $0.01. conclusion: same three scenes, same token spend – glm 5.3 flash billed $0.027 and took 36m 30s, gemini 3.7 flash billed $0.202 and took 10m 36s. glm wins gemini on price and made the best storm of the whole run follow @thehypedotnews for 24/7 ai news, analysis and breakdowns
5
403