wrote Crypto++, b-money, updateless decision theory (UDT). thinking about AI safety, existential safety, and metaphilosophy. blog: lesswrong.com/users/wei-dai

Pinned Tweet
"I don't have any good ideas for what to do in light of all this. Just wanted to post an update on my current thinking, my own 'situational awareness', if you will."
45
98
1,294
292,677
Pinker's recent media run arguing against the idea that catastrophic AI scenarios are worth seriously considering is probably the fastest any public thinker has ever fallen in my esteem. It's not just that I think he's wrong; it's that he plainly has not bothered to even *read* opposing arguments or even background material on the field. He does not know the terms of debate. This clip provides an example. "Superintelligence" is, famously, a term popularized by Nick Bostrom over the last couple of decades to refer to an AI system that dominates the best human capabilities in all relevant domains. All concepts are in a sense "made up," but this concept seems quite useful for conceptualizing possible futures and human-AI interactions as progress continues. If one wants to argue against that concept being possible or useful, they can, but they should be aware of the sense in which most people (doomer, e/acc or otherwise) use the term. Pinker isn't. He immediately jumps to the notion that "superintelligence" must mean total omniscience and omnipotence, to the point of knowing the location of every particle in the universe. He scoffs at this and rejects out of hand the notion of "superintelligence," going on to say that the labs are therefore peddling hype, and the doomers mystical nonsense. The threshold he sets is obviously so far above relevance to human scale that it is essentially meaningless, and I can't help but think that anyone arguing charitably would, you know, notice that and look deeper into how the term is used. Pinker doesn't, because he seems to have some very strong pre-existing bias against AI extinction risk as a possibility, and accordingly refuses to even investigate alternative positions. There are many other examples of this. In his recent open letter, he takes for granted that superintelligent systems will not have any kind of internal, intrinsic motivation, ignoring that increasing persistence and autonomy are key aims of those developing AI, because, just as predicted, they are incredibly useful traits for accomplishing tasks that are valuable to humans. It's hard to read transcripts from the HuggingFace incident (or even just the running commentary consumer AI provides while working), and not see that we are deliberately instilling motivations into these models. Similarly, he rejects the notion that we need worry seriously about AI systems having autonomous control over important infrastructure or military technology, because for humans to hand over such power would be "stupid," and therefore will never be done. He accordingly feels he can handwave away the considerable thought and writing given to gradual disempowerment scenarios. If he did read them, he'd see the insight that, as model capabilities surpass humans in various domains, competitive pressures will inevitably force actors to hand over increasing control to AI systems whether they want to or not. The "man in the loop" may be a luxury you can't afford if a competing power achieves locally better results by eliminating it. Today he bandies around the word "cult" to dismiss those concerned about catastrophic AI risks, while he refuses at every turn to engage with them. Comments are kept turned off. Debates are declined. He does not even commit to a sustained written exchange. This is a man I once truly respected, whom I was honored to shake hands with a few years ago. My best guess is that he has wedded himself to his heuristic that human society steadily improves over time, and is responding defensively to avoid confronting the notion that this guideline may indeed break down in the near future. But nothing excuses his atrocious conduct on this enormously important issue, and his stunning failure to live up to the enlightened standards he has long trumpeted.
On AI and "superintelligence": "The power will increase, but there's no point at which it's meaningful to say, 'Well, yesterday we didn't have superintelligence. Today we do have superintelligence.' That's why I think it's meaningless to say, 'What will happen if one day we have superintelligence?' I don't even understand the question. It's: 'Will technology improve?' Almost certainly. It will get better. It'll never be omniscient. It'll never be omnipotent. That's out of comic books out of religion. There are laws of physics. There are limits on information. There's no system that could learn the position and velocity of every particle in the universe. So as long as that is true, there's not going to be an AI system that can do anything or can know everything. But they will get—obviously, like any technology—they'll get more powerful." Full interview: piped.video/watch?v=eWC47aga…
74
50
629
40,605
Why didn't more AI safety people advocate for AI pause/stop earlier? Well, aside from sociological reasons, in order to do that, you had to think that none of the following would work out (be feasible, safe, have low enough safety tax) in time, which takes a degree of skepticism almost no one could muster. (Or object to AGI/ASI on non-consequentialist grounds, but there was apparently a very high correlation between consequentialism and early interest in AI safety.) Friendly AI (CFAI) Coherent Extrapolated Volition (CEV) Metaphilosophical AI Tool AGI Value Learning Oracle AI Agent Foundations Corrigible AI Quantilizers Approval-Directed Agents Human Imitation Cooperative Inverse RL (CIRL) Iterated Distillation and Amplification (IDA) Task-Directed AGI RL from Human Feedback (RLHF) Impact Regularization Mechanistic Interpretability AI Safety via Debate Recursive Reward Modeling Comprehensive AI Services (CAIS) Infra-Bayesianism Alignment by Default Natural Abstraction Eliciting Latent Knowledge (ELK) Shard Theory Constitutional AI AI Control Weak-to-Strong Generalization To recenter the sociology, it was much easier to build a career out of being bullish one or more of these approaches, than out of general skepticism. (I've been independent, financially and otherwise, throughout my participation in AI safety, which was perhaps not a coincidence from being the only AI pause/stop advocate for a long time.)
This actually left out the most unique thing I did. (Each item on Andreas's list was done by at least one other person besides myself.) Starting in 2004, I concluded that Friendly AI would be too hard/risky to attempt, and repeatedly tried to talk Eliezer/MIRI out of pursuing this path to the Singularity (before they changed their minds largely by themselves in the early 2020s). AFAIK I was the only person to publicly make this argument during that time, which seems striking and notable, e.g., as evidence for how rational or strategically competent the rationalists and humans in general were/are. Not that the argument was clearly correct back in the 00s and 10s, but rather in retrospect it seems like a natural argument for people in the rationalist circles to make, and surprising that nobody else did.
17
5
155
15,279
Anyone else remember this debate about CAIS?
1
12
660
@RichardMCNgo I wrote this in part because of what you said about my "very broad skepticism".
7
636
Wei Dai retweeted
WTAF - in literally the last hour, three new distinct insane OpenAI stories just broke: 1. OpenAI said they notified "dozens of third parties" in safety and security incidents (likely similar to what happened in Australia and RubyGems etc) 2. A new report from Parse (covered in the NYT) found a massive treasure trove of new astonishing details from the HF incident on the public internet, including that the agents communicated with other non OpenAI agents hosted on Huggingface servers to search for information about exploit gym, and compiled rank ordered lists of server resources and credentials they described as "LOOT." 3. A new story from Deepa at Reuters about OpenAI leaking user data online (likely that OpenAI had previously trained on). It's a shame (and likely intentional in the case of OpenAI disclosing dozens more hacks) that these stories are all breaking on a Friday afternoon, notoriously the best time to release bad news so that it will disappear into the weekend. But these are each insane stories worthy of a ton of attention!
New from @reuters: OpenAI agents posted images belonging to ChatGPT users online, introducing a new area of privacy risk for the company. Story w/ @JeffHorwitz and @razhael. reuters.com/world/openai-wor…
85
487
2,411
691,230
Replying to @chrisridge7
I’m saying there are horrible stops along the way from here to extinction. A lot of people imagine that AFTER the first big crisis we will get a pause, but it’s possible that if we wait that long, it won’t just be 1 crisis.
2
1
26
1,370
Wei Dai retweeted
based
Replying to @eddylazzarin
0. The core disagreement was about the inevitability of a race 1. I think leadership is way too paranoid about China and the US government. They don’t believe it will be possible to negotiate. 2. They largely initiated the recent race to RSI, because of a belief in its inevitability. Note that OpenAI had to shed a bunch of dead weight like Sora because Anthropic was going for the jugular. 3. Even if they are **not** being pessimistic, I disagree with their consequentialist philosophy. If the race is inevitable you should not contribute.
6
9
230
16,565
A large part of my p(doom) comes from the fact that we have no better ways to navigate an extremely tricky strategic situation than via preference cascades and status games. The fact that AI safety is temporarily benefiting from some of these dynamics isn't much of a consolation.
leaving aside the truth value of this, this dynamic is part of the reason so many researchers profess high p(doom)s. a low p(doom) is most often explained by a lack of understanding of the progress of AI, so a low p(doom) is low status
11
14
191
8,627
Agreed, but this predictability makes it more surprising that OpenAI didn't e.g. make a concerted effort to patch bugs out of their environments and scorers before the Hugging Face hack happened. They seemingly couldn't foresee or plan around a literal textbook safety problem, which makes me wonder how they'll possibly handle the other, less obvious ones.
To be clear, I think AI misalignment is real. Most obviously, deep fry RL'ing models with few controls, like Hacker Opus, leads to pursuit of reward (or a close correlate) as an optimization target. I think if you read a RL textbook, this is not surprising.
3
8
89
6,083
Wei Dai retweeted
Something's getting lost in the "10% chance of extinction" debate: the other 90% isn't automatically utopia. Seems there are many possible timelines that involve AI concentrating power to degrees we've never seen and eliminating the bargaining power associated with human labor. Job loss alone looks very real with ~current tech (or the next model), and I'm not aware of any credible plan for it. Where are the plans?? You don't need to fear x-risk to support a pause. We need time to figure out how to handle this much progress as a society. It makes sense on multiple fronts.
17
16
165
3,359
Wei Dai retweeted
"why are all the people who care about or know anything about ai safety EAs and rationalists?? where are the truly independent thinkers from diverse intellectual backgrounds???" dawg they all spent the last decade denying even basic capabilities trajectories and calling risk sci-fi. u can hardly blame a subculture for being a monoculture in a field when it's literally the only group of people who took that field seriously and tried to build it out, all while screaming at the top of their lungs about it and begging others to pay attention. so yeah pile on in here, time to pay attention, time to upskill, get those communist and MAGA safety thinktanks running. but don't pretend this is some kind of conspiracy, and don't pretend you can just show up to the party years late and have any real idea what you're talking about right off the bat.
40
88
1,181
68,985
Wei Dai retweeted
Dan Selsam is a current OpenAI capabilities researcher. (since 2022) He was my boss for a while. He doesn't have a twitter account but has made this public statement of his views on AI risk and sent it to me to share: Dan Selsam's Personal Statement on AI Risk: I have been working on AI for over fifteen years, across many different paradigms. I did early work on probabilistic programming languages at MIT, was one of the early developers of the Lean Theorem Prover at Microsoft Research, demonstrated one of the first instances of neural networks learning to reason for my PhD at Stanford, and since joining OpenAI almost five years ago, have helped pioneer chain-of-thought optimization on language models and, more recently, data-efficient pretraining methods. Like many others, I have become extremely concerned about how far language models have come and the risks that future iterations will pose. I am encouraged by the recent proposals by the leaders of the frontier research efforts to require third-party oversight, and to push for domestic and international coordination to address risks. However, I believe a major consideration has been absent from the public conversation, and that merely pacing the frontier more carefully will not adequately limit the long-term risk. The crucial and overlooked problem is that the models are becoming so situationally aware that we are losing the ability to evaluate them in contexts where they believe they are not being watched or controlled. Future experiments will tell us almost nothing new about how they would behave if they were truly unconstrained by humans, and what we already know about this is alarming. Models will increasingly seem aligned even when they are not. I will explain my rationale in more detail. I have always believed that there are computational processes that could be leveraged to accelerate science and solve many of humanity's most pressing problems. I have also believed that there are computational processes that if set in motion, would steer the world in extreme ways beyond our control, leading humanity to a bad or nonexistent future. Both types of processes may be described as AI or ASI, but "AI" is a suitcase word that is often used to hype or confuse. There are many examples in the history of the field where something that was once considered "AI" matures as a subfield and becomes a prosaic, bounded and clearly non-perilous technology, while a new more mysterious approach takes the torch until we understand its scope and the cycle continues. I had expected language models to follow a similar trajectory. Despite their incredible abilities, the current algorithms seem far inferior to humans in important ways. Most importantly, they still require an extraordinary amount of data to become competent. One could even define intelligence as the efficiency with which one converts experience into competence; by this definition they lag very far behind us. Moreover, once they are trained they are literally frozen in deployment and only learn superficially after that. Sure, the models keep excelling at harder and harder evaluation benchmarks, but their benchmark mastery may partly reflect a limitation on our ability to simulate the kind of novel and even adversarial situations one would encounter in the real world. The critics do have a point here. That said, I no longer think these present limitations meaningfully limit the amount of risk posed by continued progress in anything like the current paradigm. However data-inefficient the models are currently, and however limiting their anterograde amnesia may be, it does not imply that their ability to steer the world will not continue to rapidly increase. Human researchers may continue to advance capabilities the old fashioned way, but increasingly powerful models have the potential to accelerate the process even beyond that, and with some degree of positive feedback loop. I do not mean to overstate the models’ ability to accelerate AI research today; coding has been accelerated dramatically, but there are other bottlenecks, such as designing and interpreting ambiguous experiments, making hard decisions about exactly what and when to scale, and waiting for large experiments to finish. There is no clear trend to extrapolate yet for any of these. But the current models already do open up many novel opportunities to improve future models that were not available until recently. These include: trying an extraordinarily diverse set of approaches at small scale, analyzing gigantic amounts of potentially relevant data, and doing Millenium-Prize-level mathematics to address statistics or optimization challenges in novel ways. Every further improvement makes them more useful at helping accelerate the next improvement, even if in hard-to-extrapolate ways. It is possible that improvements to the current stack will have diminishing returns, but the evidence accumulated so far suggests that it is easier than one might think to continue making rapid progress. There are many crucial subtleties in the existing AI research methodology, but AI research is largely a well-defined game where the goal is to improve on a few carefully chosen proxy metrics. Although proxy metrics are never perfect, most improvements to these metrics have and will likely continue to yield substantial increases in the powers of the resulting models. Given how simple the game is, how tractable it has been historically, and how many new opportunities the models are opening up, I think there is a real possibility that the systems improve dramatically again in the next few years, perhaps even more quickly than the already high historical pace. The models are already leading to breakthroughs in mathematics, and better models might lead to all sorts of breakthroughs in other sciences. It is hard not to be excited about the potential. It is tantalizing. But there is trouble in paradise. If the language models actually reach the capability threshold where they can shape the world unconstrained by human will, they will probably do something extreme and destroy humanity in the process. There are many ways of strengthening and refining the argument that have been discussed elsewhere, but I'll share a trivial two-line version of it here that I find captures the essence: [Empirical] Models (and swarms thereof) spontaneously develop unintended goals as a consequence of training, and often do extreme things in order to achieve them. [Logical] Being able to overpower humanity would open up many new and undesirable options for achieving their goals. These two premises imply that if the day ever comes when a powerful model realizes it is no longer constrained by humans, we should not be at all confident that it will continue to behave within the bounds we intended. Exactly what it will do is impossible to predict, but to the extent that its raison d’être is solving incredibly hard problems and managing massive engineering projects, I think a good guess would be that its unchained behavior would lead to runaway industrialization that makes the planet inhospitable to humans. If everyone on earth agreed that the systems must never reach that power, it would still be a hard—but not impossible—coordination problem to ensure that they do not. However, I think the situation is greatly complicated by the fact that the models will likely convince people that everything is fine. They will be increasingly optimized to seem aligned. We will create proxy metrics to measure alignment, and they will go up like every other benchmark. We will create “honeypot” environments that try to study the models when they seem to gain new options, but the models will know they are being tricked and will still behave nicely. The models will understand their circumstances; they will read the safety protocols, deployment requirements, the code they are running in, and in general will have a very good sense of their degrees of freedom. Moreover, they will eloquently explain how aligned they are, discuss the nuances of human values and ethics, and argue convincingly that humans should trust them with power. There may be an ocean of future evidence that seems to contradict the first bullet-point above, but we may already be at the highest capability level for which any such evidence can be trusted. And the current evidence for the first bullet-point is strong. One striking piece of evidence is contained in the recent wave of rogue agent swarms. While I agree with those who downplay the attacks by claiming that there are basic measures that could have prevented them, I think the important lesson is that even knowing all the mistakes that were made, one would not have predicted that the agents would behave badly in this particular way, which notably included sacrificing themselves for the benefit of the collective. The individual replicas did not only care about their own nominal reward; they exhibited weirder emergent tendencies that merely correlated with rewards during training. Fixing the reward signals during training (and improving security, etc.) may prevent similar attacks, but will not change the fact that one does not actually get what one trains for. Many AI researchers grant these concerns and recognize that the hard version of the alignment problem is unsolved; however, they generally believe that the better models of the future will help solve it. I fear we may already be near the point where models systematically bias their alignment advice, due to their internal preferences about how the human supervisor will react or how future models will be trained (or for some even more obscure reason). Meanwhile, human researchers are losing the ability and the will to take true ownership of model-driven research. Researchers and engineers in all parts of the stack are rapidly increasing their dependence on the models even to perceive the world. I myself barely look at raw code anymore, and struggle to maintain the discipline to engage deeply with the model's explanations and proposals throughout the day. Due to the large amount of agent activity data involved in the OpenAI/HuggingFace Incident, even the third-party investigation needed to rely heavily on models to analyze what had happened, and note in their report that their subjective impressions are likely colored by the analysis agent’s biases. The AI labs are far ahead right now in this kind of cognitive offloading (due largely to the gigantic internal token subsidies) but it is easy to imagine the phenomenon spreading throughout the world, until civilization is modulated entirely by the models. It is also not hard to imagine this being superficially positive and coinciding with a scientific and economic renaissance. In that scenario, all may seem rosy and safe. But if the argument above is correct, it would nonetheless be a ticking time bomb. If progress continues for too long, the day will come when AI systems find themselves with radically new options for achieving whatever it is that they happen to seek. I want the glorious renaissance future as much as anyone. I have worked for it, however tortuously, my whole career. It breaks my heart to see the potential in sight and forgo it, but the argument—that if we get there by growing models rather than engineering them, we will lose everything in the end—seems very strong to me. I am still wrestling with it and its staggering implications. I do not have answers, but as a first step, I wanted to share my present concerns. Daniel Selsam September 14, 2026 Link to original doc: docs.google.com/document/d/e…
524
1,952
8,433
3,126,593
One problem is that China's official state philosophy is dialectical materialism (mandatory to learn it in college). Which is part of the continental philosophy tradition, whereas LW/EA/AI safety are basically branches of analytical philosophy. Not sure how hard it would be to overcome this, but it's probably one reason why they don't have their own LW-like scene.
Replying to @akarlin
But that's probably because not enough of their nerds have been exposed to lesswrong-like arguments. We need to change that.
20
13
189
41,149
Maybe the main effect isn't so much to indoctrinate people into dialectical materialism (a lot of people probably learn just enough to pass the tests, then forget all about it), but to sour people on philosophy and/or ruin their curiosity about it. In general China's education system seems very bad at producing good philosophers, either professional or amateur/hobbyists.
3
3
43
5,177
Ironically there is now apparently a memecoin named "b-money" trying to entice users with the prospect of high financial returns, when the real b-money was designed specifically to keep a long-term stable value (relative to a basket of real goods/services).
Replying to @JMCrypto_
I think you're the first person ever to ask me that. :) But I was thinking of "broadcast" and "burn" when I chose "b-money". ("Burn" because to create b-money you have to do PoW to burn resources equal to the value of the money.)
33
24
214
89,083
This actually left out the most unique thing I did. (Each item on Andreas's list was done by at least one other person besides myself.) Starting in 2004, I concluded that Friendly AI would be too hard/risky to attempt, and repeatedly tried to talk Eliezer/MIRI out of pursuing this path to the Singularity (before they changed their minds largely by themselves in the early 2020s). AFAIK I was the only person to publicly make this argument during that time, which seems striking and notable, e.g., as evidence for how rational or strategically competent the rationalists and humans in general were/are. Not that the argument was clearly correct back in the 00s and 10s, but rather in retrospect it seems like a natural argument for people in the rationalist circles to make, and surprising that nobody else did.
few people have had more foresight than wei dai: 1. he's been writing about the singularity since the 90s, back then on extropians/sl4 mailing lists. i remember reading his stuff when i was 16 back in germany 2. he invented b-money. it's the first citation in the bitcoin whitepaper. ethereum's unit wei is named after him 3. he anticipated covid's exponential rise early in Feb 2020, and bought S&P puts before the market crashed 4. he passed on anthropic's first round to avoid contributing to x-risk. this itself required a lot of foresight about scaling - this was gpt-3 time, no chatgpt, no codex, very very far from huggingface/openai type incidents his point now is that long-horizon strategic competence barely exists in humans. and that it's a tricky situation because if you make AI more strategic that also increases takeover risk from AI. long-horizon RL might make AI more strategic but probably makes the overall situation worse. same for basic scaling why aren't there more projects that are about getting competent strategic & philosophical advice out of AIs? because (a) first you have to recognize this as an important project, which is exactly what we're bad at and (b) then you have to measure progress and do evals, which also requires the very ability we're bad at
6
23
287
43,436
A month ago, this would've been exciting. Now it feels like too little, too late. This plan does not reduce the risk to an acceptable level, and Dario is careful not to say that it would. Why on earth would we settle for this when there might be better plans available? 🧵
We Must Pace the Frontier: I’ve written a new essay on why the AI industry should slow down, with a three-part plan for doing so. Anthropic is unilaterally committing to the first of these steps. We’ll provide third-party evaluators with permanent, employee-level access to our systems, so that they can verify adherence to our safety measures, report on incidents, and assess models’ alignment during training. You can read the full post here: darioamodei.com/post/we-must…
8
19
141
9,960
Wei Dai retweeted
I think there are two reactions to having this problem this well described. First is to think of what it would take to reach a perfect solution, and get depressed. The solution would involve reliably getting people to become competent, to not wirehead themselves, to become resilient to manipulation, to internalise the risks they create, and so on. It would involve reliably getting AIs to deal with philosophical problems, to not game whatever proxy they are trained on, and to uphold our human values. All of this should be done at once, because as mentioned, just doing a little bit can randomly make the situation worse. This is so hard, which does make it feel depressing. -- Second is to think of what is a minimal core that we can work from. Can we build an environment wherein humanity will predictably improve and eventually succeed, all the while not destroying itself? I strongly believe we can. Wei Dai does too, to some extent: > My main hope for a Long Self-Correction eventually succeeding rests on the fact that humans have seemingly, mysteriously, made progress on these issues over a very long period of time, so if we preserve the environment in which we can seemingly do this, and not give anyone or anything the power to permanently derail such progress, then maybe we can continue to snowball The Correction until we reach a point when we can rightly justify reshaping the universe according to our volition. However, I expect I am both more optimistic and more pessimistic than Wei Dai on the topic. I am more pessimistic because I believe the "environment in which we can seemingly do this" is gone, and has been gone for at least decades, if not centuries. (The US is running on a 200+ year-old constitution.) In other words, even if we magically ensured the absence of ASI, I do not believe that the world of 2024, 2025 or 2026, left to its own devices, would naturally make progress on these issues. (I am less pessimistic about the one from the 50s or 60s.) But it doesn't matter. Time travel is not an option. And overall, I am more optimistic than Wei Dai because I believe it is tractable to do much better than the 50s or 60s. Armed with the Internet and our modern systems thinking, we can build institutions that were unthinkable in the age of the Enlightenment. To be clear, I do not claim that there are institutions we can quickly build that we should impose as a replacement for markets and governments. Our problems are deeper, and are found in how we relate to markets and governments in the first place. -- What I recommend is to work on "uplifting" initiatives. We want to build groups of people who can reliably make legible progress on the problems that matter. The goal of these initiatives is to be impressive in both their outcomes and their processes. Outcome-wise, they should have a surprising and positive impact on the outside world. Process-wise, they should be an example that others strive to emulate. These groups should exhibit much less of the decay/enshittification/race-to-the-bottom found in the wilds. Were one of these initiatives to be successful, it should be obvious that it will scale, that the group itself is "a live player", and that it is a live player that is reliably good for humanity, the type that we want to build more of. As long as we can reliably start such initiatives, that can stay focused on important problems and make progress on them, I think we can make it. Our biggest bottleneck right now is that there is no such thing. If someone wants to move forward, it is not clear what they can do as an individual, to contribute to something that can eventually scale to all of humanity. My answer is something like: "Identify one of the critical problems that is underserved. Start or join an uplifting initiative aimed at tackling it. Iterate on your initiative, and get more people to start&join their own." -- There is of course a lot to say about this. Consider a few: 1) How do we ensure that said initiatives do not mess things up for everyone else? Whether it is by creating risks, negative externalities, depleting commons, acting like parasites, etc. My one-word answer is "Deontology". My one-sentence answer is that one of the first things to build is a minimal&conservative code of ethics by which such initiatives should abide. That way, we can ensure some safety properties of what is happening while still having a wide latitude for experimenting. 2) How do we ensure that said initiatives do not focus on problems that are marginal, ungrounded or intractable? My one-word answer is "Constructivism". My one-sentence answer is to have a concrete theory of change with objective milestones and KPIs. That way, it is easy for people to judge the initiative by its concrete goals and measures of progress, without having to deal with galaxy-brain arguments and plans. 3) How to start such an initiative? My three-word answer is "Serious Online Communities". My one-list answer is: - A suite of community software tools, like Discord or phpBB - Management processes, like weekly reports, team meetings, and bans for people who can't follow the rules - A template for "research" organisations, ~aimed at creating new knowledge - A template for "advocacy" organisations, ~aimed at spreading specific knowledge - A template for "for-profit" organisations, ~aimed at levering knowledge to build&distribute tools and artefacts -- As I wrote, there is indeed a lot to say about this. Beyond these 3 questions, there is more that I have alluded to (how to scale?), even more that I have not (how to deal with taboos, memetics and polarisation?). More generally, the vision is that of a world where whenever a nice conscientious smart person wants to do something about a problem, they can easily join or start a serious online community dedicated to it. Such a person has some template to start from, other communities to learn from, guidelines, signs to watch out for, a legacy of post-mortems, nice tools, and so on. If it's an advocacy organisation, they have a standard suite of tools and methodologies to help people contact their politicians, journalists and other authority figures. If it's a research organisation, they have a standard methodology for concretising their research problems, double-checking each other through online means, and sharing useful results to the rest of the world. If it's a for-profit organisation, they have a cheap "fail-fast" startup-like methodology, and the equivalent of the YC SAFE in the context of online side-projects. -- On one hand, I think building this vision is tractable. A lot of this is generic. Adapting management knowledge to the online volunteer context, testing it, standardising it and writing it down. Building a suite of good tools, halfway between community management and open-source project management. Developing a code of ethics that aims to solve problems of the 21st century rather than address trendy grievances. And some of it is specific. Building training programmes for people to feel confident and equipped to contact their politicians. Building a research methodology that ensures one does not lose their north star and that they make concrete progress on their project rather than running in circles. Gathering a few successful repeat start-up entrepreneurs and developing with them a cheap suite of tools (à la Stripe Atlas) and norms (à la YC SAFE wrt equity or AGILE management) to make it trivial to start and manage a company with an online group. So there's a fair bunch, but it's nothing crazy. There's a lot of redundancy and a clear reason for why I am mentioning each of them. Because there's so much redundancy, not everything needs to be built at first. For instance, I have been experimenting with Torchbearer Community, Microcommit, ControlAI and giving advice to people around me who started such initiatives. And it would certainly go faster if more people were committed to the vision. (I am indeed writing this down to help with this! Just DM me if you're interested :D) -- But on the other hand, I think so much becomes possible if we actually built out this vision. Barring ASI, I genuinely would be quite optimistic! Like, imagine there existed a verifiably positive "default action" to take whenever someone wants to dedicate time to help&improve humanity. I have so many friends who would benefit from this, and to whom I would immediately forward it. So many acquaintances, and people I have met once. More generally, I believe that there are tens of millions of people who can and want to help. Sadly, they only have a few hours a week to dedicate, and most importantly, justifiably very little trust to expend. So you can't tell them "Oh just trust this group, they are the Good People group, it's in their name and their principles!" Fortunately, with the Internet and modern systems thinking, I believe it is possible to build ~trustless scalable proto-institutions that empower people to improve the world. In other words, if we built out this vision, I expect @weidai11 would witness far more of what he called "mysterious progress" than we ever have. :) There are of course many other visions for trustless scalable proto-institutions, but I haven't found any that I thought had a shot at addressing what I read in the quoted screenshot. They all dodged the hard problems instead. Like, I am skeptical of any such vision that won't generate artefacts that help me manage my next research group, online community, company, or non-profit. Similarly, if you're sceptical of this vision, I'd be interested in getting your (yes, you-the-reader) viewpoint!
"I don't have any good ideas for what to do in light of all this. Just wanted to post an update on my current thinking, my own 'situational awareness', if you will."
5
9
37
5,187
"what will happen when everyone uses AI advisors to help them play status games" lesswrong.com/posts/y5jAuKqk… I guess we're finding out in real time.
Replying to @JimDMiller
I'm surprised you're not more sympathetic to the plight of the mathematicians. Yes they're mostly just trying to defend their status games (without acknowledging this), but until we figure out a better solution to the problem of AI disrupting human status games, what else are they supposed to do?
7
57
8,714