Standard CDT agents defect against each other in prisoner's dilemmas.
@allTheYud once told me that his research on FDT was driven by the deep-rooted intuition that superintelligences will surely find a way to do better than that.
But Eliezer's worldview implies that agents which become smarter over time are locked in adversarial relationships with their past and future selves (as they move toward alien squiggle-maximization-like values). My own research is driven by the deep-rooted intuition that superintelligences will surely find a way to do better than that.
I've cared about this on a personal level since childhood, but I didn't realize how foundational this idea was for my broader worldview until I reread my sci-fi collection. In one way or another, most of the stories in it are about how to remain loyal to your past self, and how to trust your future self. And even my political worldview is driven by the idea that there *must* be a way to establish healthy relationships of trust and loyalty between Western elites and the people they govern, despite elites typically being smart enough to outmaneuver much larger populations of common people.
I haven't yet figured out how to pin down these intuitions in a formal theory, but I do have a hazy sense of its outlines. It has something to do with agents deliberately constructing entanglements between different decisions, and cutting through the infinite recursion of me modeling you modeling me modeling..., and establishing identities with sharp boundaries and functional credit assignment protocols.
These phrases are not intended to be clearly interpretable to almost anyone, but it feels important to me to make this quest at least partially legible, so that others can start to notice and nurture this same intuition when it arises in them.