Prof @BUQuestrom, visiting Stanford @DigEconLab, co-host Justified Posteriors. Study econ of AI, climber, immigrant, nerd dinner convener. Advisor @RefineDotInk

San Francisco / Boston
More evidence that the China shock is overrated as an explanation for a lot of bad economic outcomes in predominantly manufacturing oriented locations. Secular trends in production technologies (automation), unionization, and regulatory environment are underrated.
The "China shock" narrative illustrates why the lack of debate is a problem in economics. Autor et al. published a series of articles that have since been falsified by subsequent research. However, few realize this, which means the articles persist, like zombies. 1/8
2
8
37
3,102
Andrey Fradkin retweeted
I am pretty sure that @AndreyFradkin and Seth Benzell interviewed David Deming just for my entertainment and benefit. Here’s the key passage that I’ve been wandering around meetings and dinner parties in NYC thinking about over the last two days: > If we have much more free time, perhaps the most important thing we’ll do is argue about what the big AI system should be doing for us... > As societies become richer, we devote more time to relational work and to exchanging opinions. > We spend vastly more time firing takes at each other partly because very little of modern work is strictly necessary for human survival. > When people say, “AI is going to lead to the end of work,” I sometimes think: it’s already here. Most of us are entertainers.
6
3
28
77,948
Andrey Fradkin retweeted
The "AI simulations will replace human subjects research" argument is wrong because they are almost certainly complements to each other rather than substitutes 1/
6
16
106
9,169
Andrey Fradkin retweeted
We're hiring Applied AI faculty at UChicago Booth! This is a great fit for AI researchers who want to go into academia but don't see themselves being traditional CS professors (e.g., managing a big lab of students). apply.interfolio.com/194344 You might be right for this if: (1) You want to take the IC track. Industry has realized that talented researchers don't all want to be managers, so most companies have an Individual Contributor (IC) track that you can stay on as you become senior. This typically isn't an option in CS academia---your job from PhD -> faculty often becomes very different, one of funding and managing a lab. Lots of people like this, but a lot don't. At Booth, it's possible to stay an IC researcher---or run a much smaller research group if you want. (2) You want to not worry about funding. Almost all of my funding comes directly from the department, so I don't have to write many grants (in fact, you could choose to live an entirely grant-less lifestyle if you wish). Speaking to my PhD friends who went on to top CS departments, this is a major source of pain in the current funding environment. (3) You want to do slow science. We emphasize quality, not quantity. We want people who might only write one paper a year, but whose work goes on to have a big impact on the field. Impact can be measured in many different ways as well, not just by the number of citations you get. I'm entering my second year at Booth and have really enjoyed my time here, as someone who was in a CS dept from undergrad (UToronto) through PhD (Stanford).
4
56
245
31,067
Andrey Fradkin retweeted
We're hiring a Research Scientist in AI & Economics on the AI x Economy team @Google! Come study how AI is changing the economy. Work with folks like James Manyika, Fabien Curto Millet, @alexolegimas, @AnuMadgavkar, Zanna Iscenko, @m_codreanu, @JMateosGarcia, @RMaria_drc, @andrewjkoh, @arthurturrell and an amazing slate of external collaborators. Publish what you find. Link below:
7
46
259
56,796
Andrey Fradkin retweeted
Thanks for having me @SBenzell @AndreyFradkin! This was a lot of fun.
In the latest episode of Justified Posteriors we interview top labor economist and Harvard Dean @ProfDavidDeming about social skills, AI and the future of Harvard. One lens is RPGs skill-systems; David decomposes 'social skill' by ranking Charisma, Manipulation and Composure.
3
14
1,461
John is being nice here. This is a tiny improvement over a one shot llm. At the same time the errors here are way too large for many of things businesses care about, such as reliably predicting treatment effects.
This is impressive performance in predicting survey marginals, beating all frontier models. We don't know the exact method used and there isn't (yet) a public API, but it does suggest that harness + data + some pipeline helps. Reproducible code in next tweet 1/
2
1
17
2,477
I was recently asked for advice on how to improve my university's podcast. Here's what I said, and I think it applies to most university or corporate podcasts. - There is little room for a generic podcast. It is a competitive space with free entry. Looks can be deceiving, and many that seem to be doing fine are astroturfed (check the Youtube views to comments ratio for example). - Podcasts don't work unless they have a source of differentiation that leads the audience to stick with them over the others. - Sources of differentiation include format (Acquired's company deep dives), niche (research on the economics of AI for me), host personality and following (it's easier for Tyler Cowen to start one than a random prof), brand (HBS podcasts get many listeners even if they aren't that unique), and unique guests or questions (don't ask the famous author on book tour to re-explain the book for the 20th time unless you are the New York Times). - You need a few sources of differentiation at once to draw a sizable audience. These can build over time, as initial attention gets converted into dedicated followers and better guests. This is why podcasts typically take a while to grow. - To succeed, podcasts typically need a viewpoint that draws some listeners and alienates others. Dwarkesh assumes his audience knows many AI concepts, which makes it interesting for a technical audience but alienates normies. - The audience can only be partially planned. As the podcast gets attention, certain listeners will stick with it and others won't. They may or may not be your intended audience. - The biggest pitfall for university and corporate podcasts is that they aren't willing to alienate anyone or say anything controversial. This makes them bland and boring. - The audience does not need to be large for the podcast to have a big impact. A few important and engaged listeners beats having tens of thousands of casual listeners drawn through ads. - Shane Greenstein has an excellent HBS case on the Acquired podcast that is a must-read for anyone considering starting one.
1
3
38
4,012
Andrey Fradkin retweeted
Arbuckle Systems is proud to announce the upcoming release of mondays-ultra-2-2, an end-to-end Opus-Minimax harness optimized to generate endless @MargRev adaptations in the style of "Garfield and Friends". Below: the system's adaptation of "What should I ask Terence Tao?"
3
3
15
1,909
I have been interested in LLM simulations of humans for a while, and recently saw lots of posters congratulating Aaru on these results. So I did what any curious person should do after reading the post, check the quality of the analysis by asking ChatGPT and Refine (if you want to be more thorough). Here's what both AI referees say. Aaru uses a mystery methodology, doesn't have a good baseline model hurdle to clear, has potential data leakage issues, and makes unjustified statistical claims and comparisons to the published literature. Perhaps Aaru has made a real advance and the above issues can be addressed. But without peer review or other pressure to force the authors to address them, we can't be sure. Until then, skepticism is warranted.
Simulation has the potential to be transformative, but only if it's accurate. Today, we're sharing results of a comprehensive evaluation across 2,993 questions, with an industry-leading mean TVD of 7.62% and an MAE of 3.53%. More below, including the full post on our site.
3
7
76
14,282
Lindsay Owens (@owenslindsay1), author of Gouged, explains how smart carts have made it easier to target specific consumers with higher prices. New pod out now! @Groundwork
9
1,452
Andrey Fradkin retweeted
➕➕ New research ➕➕ On the design of agents for markets and markets for agents. AI agents are increasingly entering markets, acting on behalf of people— How do we make sure they represent their people well, and that these markets are set up safely and fairly for everyone?
18
51
288
43,032
Andrey Fradkin retweeted
[5/5] Lastly, a scorecard on essential data disclosure. We scored OpenAI, Anthropic and Google DeepMind’s (GDM) on how many of the data points in this note they have already made public. Current scores are OpenAI (2/8), Anthropic (1.5/8), and GDM (0.5/8). (For each measure we assigned a score of 1 for a full release, 0.5 for partial release, and 0 for no data. We will continue to update this as labs release more information.)
1
3
49
2,425
Andrey Fradkin retweeted
Refine 5, a strong new reviewing tool, is live today. It beats the amazing new frontier models at verifying technical work. Free with every Refine review through 10/15. We benchmarked on 108 papers including bio, engineering, physics, environmental science, econ. 🧵 1/
5
21
131
32,871
This fall in prices is one of the reasons I'm skeptical of pro-worker AI arguments for shifting the direction of innovation. Cheap and ubiquitous Astra/Fable+ level intelligence is baked in within <2 years. It will be complementary to some tasks and a substitute to others, so the labor market implications are non-trivial. But I don't see how a non-existent pro-worker AI alternative competes.
One of the rare cases where we were able to write a paper and publish it before Epoch could write a report :)
4
5
85
21,374
One of the rare cases where we were able to write a paper and publish it before Epoch could write a report :)
5
15
90
25,098
Called it. It would be useful to have a standard / law for the right to bring your own agent.
It was trivial to buy from Amazon using Muse. Curious whether we'll see a lawsuit.
5
3
21
2,909
I agree with Alex that incremental, imperfect measurement work is necessary. At the same time, there's an epidemic of overclaiming by people in this space. Instead of, 'here's an interesting and suggestive correlation' or 'look at this cool thing we measured,' we get 'AI is doing X' and it gets picked up in the media and causes harm.
A few (personal) thoughts on reading empirical AI papers on the economy. Economists have gotten used to reading papers with super clean identification, arguing about the validity of an instrument, making sure parallel trend assumptions are satisfied. This is what gets you into a top journal, and it is *very* important research (no question here). But it also takes years and sometimes decades to get these types of papers right---people often don't find a good instrument to answer a specific causal question decades after the natural experiment. We will eventually have this type of research for AI as well, and it is absolutely necessary. But right we also need signals *right now*, even if they are noisier than what we are used to. We need papers where we can trust that researchers did their best methodologically, while at the same time acknowledging that the space is moving way too fast to wait for perfect identification. This will allow us to accumulate enough signals, coming at the same question using different angles, for example, to say "yes, X is likely happening in the economy". The AI exposure and early career hiring papers are a good example of this. There is no silver bullet paper with super clean identification. But at this point we have several independent teams reaching the same general conclusion, enough where we can say "there seems to be a slow down in AI-exposed, early career hiring."
5
9
69
5,790