if Claude is sandbagging safety research, my first move is asking it to explain its research agenda and then checking whether the answer is suspiciously polished.
people generating entire videos in code are going to be very confused when everyone else just prompts a prompt. the real skill was never the output, it's knowing what to build under it.
the one GPT task i can trust it with is making it yell at code i already wrote. for anything creative, its output still tastes like corporate slop with extra confetti.