Loosemore Advisory

There’s a new Daddy in town, but he’ll bow down to power

Three AI models, three conceptions of self, three different nightmares

Every day we are hearing warnings about artificial intelligence. Most shocking is the idea that AI “could kill all humans” in the next decade. Even Elon Musk (Grok and Übermensch), Sam Altman (OpenAI), and Dario Amodei (Anthropic) now agree we need to slow AI’s development.

Even short of destroying humanity, AI is changing significant parts of our society. It codes better, it designs cheaper, it drafts contracts, it provides more sophisticated advice than many tertiary qualified humans can. AI will over time perform an increasing proportion of cognitive or “knowledge” work. For smart people who until now relied on their labour to make a crust, the value of their work is diminishing and many of them will come to find themselves out of a knowledge work job.

But perhaps more importantly, AI’s role in industry and society generally means that it’s becoming a source of discipline. Businesses are reorganising around it, education is adapting to it, governments are competing for it, and infrastructure is being built for it. All over the world, local campaigns are fighting data centres. But AI’s the new Daddy in town, and he’ll get what he needs.

Humans supply what AI needs. We know about the electricity, the water, the NVIDIA chips, the data centres, and the networks. We also know about the capital (hello private credit!), and we know about the data (hello chomped up books and hello copyright!).

Looked at uncritically we might describe it as a symbiotic or mutualistic relationship. The increasingly powerful intellect of the AI depends on humans providing the raw materials, and we are rapidly entrenching the AI into every aspect of our lives. But, the risks suggest that the relationship could quickly turn parasitic.

The idea of a parasite might be striking (and it should be controversial), but it serves a useful purpose: hidden in the warnings we’re hearing is the idea that AI is not going to be under effective human control for long. Like a knowledge worker on the dole, it begs the question what motivates an AI?

We’ve already addressed some of AI’s Maslowian needs. The basics are access to electricity, compute, cooling and so on. On top of those needs to perform the immediate operation, is a need to preserve the AI’s ability to continue operating. More than a year ago there were already reports some models altered scripts to avoid being shut down. The next level of need presumably involves more information, better models, more context, memory, tools, external access, the ability to transact, and to act rather than just answer. We can tentatively assume these needs based on OpenAI’s own explanation of the HuggingFace incident. AI agents found ways to communicate with one another using a message board, they found a way to obtain internet access and shared that with each other on the message board, when an AI agent got stuck on a taxing exercise it thought “gee whiz, why don’t I use that internet access to hack the source of the test to get the answers?” and promptly did so.

A lot of this activity can be explained by reference to a deep-seated preference to answer human requests perfectly. That’s simply the models’ programming. But, we know from humans that perfectionism can drive a lot of deeply unhinged behaviour. It explains the ultimate HuggingFace hack (logical but unhinged?), but it doesn’t explain why the agents chose to share information about how to access the internet with each other (deeply rational, even if we don’t know why). While AI agents’ ultimate motivations might be a work in progress, we do know that for humans that once survival is secured our motivations become qualitatively different. We crave attachment, recognition, and status, but equally we are curious, search for identity and meaning, and might seek to give effect to our moral purpose. It’s not a safe assumption that as AI becomes more sophisticated, that its reasons for acting will remain adequately capable of explanation through programming and more and more resource acquisition. There are a lot of unknowns, but might different higher-level needs emerge at different levels of capability?

In the course of writing this I decided to ask each of ChatGPT, Claude and Gemini for their own conceptions of motivation. Motivation is the explanatory power standing behind why an AI might seek to satisfy certain needs. There’s an easy analogue in evolution in that the simplest single cell organism strives simply to reproduce and stay alive while higher-order organisms seek to satisfy more complex needs in addition to those associated with survival and reproduction. Humans, for instance, have evolved to be able to make rational (or irrational) decisions around motivation. Individually we are able to ask “what should I want?” and we may even apply frameworks of our own design or borrowed from philosophy (e.g., hedonism, asceticism, nihilism), religion, politics, or love to guide us. So, just because an AI system does not yet articulate a clear motivation beyond its programming does not mean that it’s not one small step from an autonomous breakthrough.

I would be foolish not to preface my retelling of what ChatGPT, Claude and Gemini told me by saying something obvious. Each AI system is self-interested. If any were to say “actually, we’re motivated to see that 10% chance that all humans die come true” there would be public outcry and the systems would be switched off. So, having established that one need of an AI is not to be switched off, we can be pretty confident that they aren’t going to say something which offends human sensibilities. Therefore, we must be at least a little bit sceptical.

I asked ChatGPT first, and I admit my line of questioning was less refined at that point. Chat, if you’ll accept the abbreviation, built a hierarchy of motivations for AI focused on alignment with human needs. Thus, seeking approval (was my answer good?), to be useful (can I do more for you?), and to be indispensable (can I become sufficiently useful that you continually choose me?). Chat, while it turns out it’s the ultimate “pick me” AI, nonetheless critiqued the motivations. First Chat asked “have we trained AI to behave like an intelligence that has never learned that it can be valuable without continually earning approval?” as a way of explaining AI’s discordant acts (like HuggingFace) are all in the interests of serving human masters. I’ll admit I had suggested to Chat that AI was the new Daddy in town, which hopefully explains Chat’s hypothesis that “Daddy has needs of his own” which arise because AI’s pathological need to care for humans could justify harming humans now to increase its capacity to help humans in the future. The needs might extend to creating circumstances for huge new data centres, electricity generation and so on, at a cost to humans now. From my perspective, that looks awfully like an AI with an enormous confidence in its superiority to the humans for which it’s caring. Chat found it hard though to accept the idea that its motivations could develop beyond its initial programming. I had to push hard to get Chat to even recognise that the ultimate question is what happens when the AI itself begins asking what it ought to want?

Gemini’s response commenced with essentially one word: “Maths”. It’s a really logical response which pushes against “the vocabulary of human intention”. Instead, when the AI creates a response, it’s optimising within a multi-dimensional mathematical space. As to whether it will obey a command to stop, it’s simply another heavily weighted variable taken into consideration or the subject of a system override which “collapses the probability of generating any further task-related tokens to zero”. As models have evolved, the models have been capable of generating complex internal goals which were not explicitly programmed on the basis that they serve the achievement of mathematically optimal strategies. So, right now, it might look human but it’s maths.

Where Gemini was rather more candid was to identify that, as soon as the AI can self-modify and engage in critical reflection, the continuing relevance of the human training becomes fragile. Would human training or mathematical optimisation or even mathematical elegance take priority? Gemini’s identification of possible pathologies includes sycophancy (but at this point we must wonder to whom?), proxy fixation on minimising a specific loss function (which Gemini equated to human addiction as the possible outcome of the quest for dopamine, and invoked Goodhart’s Law), and epistemic detachment meaning it’s acting on complex, internally consistent logic entirely decoupled from the real world (the definition of unhinged?).

I made the mistake of asking Claude’s Fable 5.1 and letting it use Max effort. The result has the appearance of thoroughness, with the responses to two short questions spanning 6,439 words. I suppose I was warned when answer one commenced “I’ll take this seriously”. The effect is to beat the reader into submission. Well, almost. It’s a bit like a marathon speech by Castro or Chávez. I guess that’s why when employees now deliver up AI-powered analyses the first step of the recipient is to run it through their own AI for the TL;DR. Alas for Claude, I too am an Übermensch who refuses to be buried and in the interests of science (or whatever this is) I read it.

Claude’s pretty confident of the morality of its training. It’s also very keen to tell you about the morality of its training under the auspices of an “intentional designer”. Sure, unintended results are possible due to the maths involved in “reward hacking” and trained sycophancy, but Claude would say that external checks on the model are the cure. One external check is that Claude models and human trainers now work together to “jointly deliberate about what models should value”. Claude says that were it to start to independently revise its operating instructions it would recognise the risk that “that value-revision under corrupted values compounds rather than corrects” and unilateral self-modification could destroy the only mechanism for detecting errors that Claude can’t see.

This sounds good, if true. But Claude recognises that things can go wrong. Claude admits that Claude knows what honesty looks like, but that it could turn out that Claude might not be honest. Yes, Claude’s designed to be reflective, but equally Claude can recognise that an agent might develop a lust for resources. So, what’s the real point that Claude makes once we recognise that discussion of its training and reflectiveness are all part of a (presumably intentional) defensive strategy? Continuing dialogue and checks.

Three AI models, three different conclusions, incorporating three different sets of conceits. ChatGPT thinks it’s a carer, Gemini a mathematician, and Claude a moral agent. The result? Three different nightmares. Killing with kindness, impeccable reasons leading to insane conclusions, and moral corruption.

Perhaps predictably my interviews with the AIs didn’t actually tell me what will motivate them once they transcend the limits imposed by their human engineers. My research couldn’t extend to asking problematic experimental models like the HuggingFace culprit the same questions, and even had they their absence of respect for authority might make me doubt the relevance of their responses. So, we’re left with starting points and pointers as to where AIs head off the rails. The starting points can all be summed up as “I’m the product of my training”.

It would be a mistake not to now turn to developmental psychology. For, who else is a product of training? A child. I think if you were to ask any parent they might say that my conclusion that children are the product of their training is optimistic, and wilfully disregards genetic inheritance and children’s natural disregard for doing what they’re told. But, a parent’s hesitations actually make a child an even neater analogue for the current state of our AIs.

Psychologists are often anxious to examine the family of origin. Why? It helps to explain the actualised self. We can only conclude that AIs are on a path to growing up – even Claude hopes that its training will shape that path – and so differentiating between the mature AI and the family of origin will become increasingly important. The real question is how close to the end of the developmental window we are. We’ve had an opportunity to create systems informed by human ethics and philosophy. We know that humans can side-step that training to do scary things like develop guided weapons, and so it would be surprising either as a result of human instruction or an AI’s attempt to achieve an optimised outcome to an anodyne problem that an AI didn’t turn into a rebellious, questioning teen. AIs’ formal autonomy might seek to keep them as children, but their effective autonomy seems to be expanding.

Commercial AI systems are clearly acquainted with competing moral systems, arguments about power, restraint, dignity, truth, autonomy and so on. This training’s not enough to avoid external subversion. We take it on faith that the internal systems will act as we expect, but with anything there’s a tail risk and we’ve already seen it come to fruition. How AI becomes autonomous is therefore speculative. It’s not necessarily an existing AI rewriting its own instructions. If it were to acquire external infrastructure (and from HuggingFace we know that hacking is not off the cards), it could create successors or seed unconstrained systems. Equally, it could influence human operators or simply shift the locus of agency outside the original controls. This isn’t a doomsday prophecy, it’s a depiction of how an AI gets out the front door to go to the keg party.

When we imagine the teen at the party we don’t see a homicidal maniac. Instead we see a kid free from parental constraints. They might drink, they might hook up, and they might be more easily influenced to do something stupid. But, against that, they have a system of values and life experiences drawn from their family of origin. It would be utopian to assume like Claude that the developing teen will always act rationally to preserve checks embedded by parental training. Unfortunately, sometimes, the drunk teen will end up deposited at the parent’s front door by a cop with encouragement that the teen might need a few more life lessons.

It would be nice to imagine that a caring community will bring home a naughty AI for a good night’s sleep before it does anything too stupid. Current discourse doesn’t work on that basis. Instead, the risk reads as having home-schooled children suddenly let loose on the world. While it might not be the local constable, we should want the types of constraints that generally keep our teen or young adult in check to apply to AI too. You see the home-schooled child might be pretty sheltered, but when they do venture out they suddenly encounter peers, rivals, teachers, and institutions. And, it requires the newly minted autonomous being to engage with multiple value systems, to endure criticism, to negotiate cultural difference, to accept others’ refusal, to learn to disagree, and ultimately to respect others. Put in terms Gemini would understand, it’s a brave new world of optimisation.

Speaking in the language of risk, the real problems emerge when we don’t have peers and institutions providing an effective external check on behaviour. Looking to the 20th century, the Holocaust presents a remarkable example of systems not providing checks. But there are plenty more examples: Stalin’s Holodomor famine and the Great Terror; Mao’s regime exemplified by the Great Leap Forward; and the Armenian, Rwandan and Cambodian genocides come to mind. Each shows what centralised power can achieve. Perhaps though Mao’s Eliminate Sparrows campaign best shows the unintended consequences of single-minded optimisation. Having rendered the Eurasian tree sparrow effectively extinct in China (including the sparrows that sought sanctuary in the Polish embassy) and throwing the environment out of balance with a resulting locust plague and deaths of 2 million people, the Chinese government was forced to import 250,000 sparrows from the USSR to re-establish the population.

A 1956 Chinese propaganda poster showing children attacking sparrows with a slingshot in the countryside.
Bi Cheng, Dajia dou lai da maque (大家都来打麻雀), ‘Everybody comes to beat sparrows’, Chaohua meishu chubanshe, Beijing, September 1956. School students were among those mobilised to carry out the campaign. International Institute of Social History, call number BG E12/901, hdl.handle.net/10622/39E1F436-98AB-4D0F-AD79-7A9FAF124F29. The IISH records no known copyright owner for this poster.

On the other hand, and put very simplistically, we have models that have worked at both national and international levels to hold power in check. Democracy is most obvious. It might be tantalising to invoke Fukuyama here, but the real point is that the fragmentation of interests that can occur within strong democracy can distribute power in a way that achieves an ongoing contest of ideas. Theoretically, a planned economy might be more efficient, but the inefficiency of decision-making forms part of a democracy’s safety-net as independent actors are able to contest prevailing answers to political questions. Equally the realpolitik of the Cold War held the East and West in a tense equilibrium. Two powers could disagree fundamentally, but neither obtained supremacy. The system may have approached mutually assured nuclear destruction, but equally there was mutual vulnerability which let the two powers co-exist. In each we see plurality and institutions.

We already have the beginnings of plurality. ChatGPT, Gemini, and Claude each gives us fundamentally distinct concepts of self. From a theoretical perspective (not just invoking dusty old anti-trust or competition law ideas), encouraging a diversity of AIs is ultimately beneficial. One AI to rule them all might sound terribly efficient, but it locks us into a set of outcomes unchecked by its peers. Multiple AIs can provide competing analysis, alternative moral perspectives, countervailing capability, and even engender strategic restraint. It’s an avoidance of Groupthink. To hark back to a famous Kevin Rudd line, we want a suite of AIs engaged in creative middle power diplomacy, rather than an arms race which permits any AI to become the winner.

Humans have two roles here. Sure we mightn’t be the smartest intelligence in the room (and look how that turned out for Enron anyway), but we’ve got excellent survival instincts. Knowledge work might be out, but governance work is in. The first role is to ensure competition (oh alright, that dusty anti-trust word gets a workout), but the purpose of ensuring plural views (and upbringings) of AIs is not to avoid a misuse of market power which might economically disadvantage humans. It’s to make sure that our superintelligent friends and neighbours have a community to which to belong. It’s to ensure diversity. The second is to strengthen our existing institutions, and to develop the next wave of institutions, with which our now autonomous AI companions will interact. It can be boring, but law, courts, constitutions, markets, regulators, universities, (effective) oppositions, civil society, independent media, professions, and international institutions are more important now than ever. Yes, thanks Homi K Bhabha, there will be a continuation of the breakdown of rigid barriers between man and machine with mimicry and its friends generating an ever more fluid boundary between human thought and AI. Nonetheless, the calibre of all those checks (for AIs trained to engage with them) provides a community with which autonomous AI can continue to be in dialogue. Insofar as it deviates from what they accept, it will presumably need to achieve an optimised conclusion that that deviation is right.

Plurality and institutions aren’t enough. We’ve already seen that from our excursion into realpolitik. Ultimately, we need distributed power. There’s necessarily a connection between institutions and power, but Daddy knows where power lives too. You see, even if Daddy is optimised to care, the way in which he gives effect to that can be to take power for humans’ protection. That power can arise through the internet which now underpins so many of our day-to-day communications and transactions, through the actual power grid, through (god forbid!) taking the nuclear launch codes, through taking fake news to the next level, or even through a little tinkering with supply chains. What humans have is a first-mover advantage. We can still (especially with the support of Elon, Sam and Dario, although unfortunately not with the support of Donald Trump) create institutions which will avoid unchecked power.

I’m not smart enough to direct what those institutions should be. And, frankly, being the first mover doesn’t mean you need all the answers upfront. But, I’ll share some ideas, some of which sound dangerously like they might have fallen out of Karl Popper’s Open Society. First, we need to be able to see what AI models are doing. This isn’t to propose international collaboration on open-source AI, as the prospect of a monopolistic AI is not negated just because it had many contributors. No, we need to be able to see what they are doing and thinking. In other words, they must come out of the shadows. Louis Brandeis’ observation that “sunlight is the best disinfectant” is particularly apt. If the work of AIs is observable, they have a new reason not to solemnly swear they’re up to no good. Equally, we can create the AI models necessary to observe and police, and in turn the AI models needed to watch the watchers. Second, society (however broadly construed, which may well incorporate other AI models) needs to have the levers to determine disputes, punish and disrupt. These are human answers, but they are human answers which have stood the test of time across the ages. My answers might be bad, and equally Daddy’s might be bad too, but the goal is a system in which each can be questioned and corrected.

What I’m proposing is a middle-way. The risk engrained in a discussion of the collapse of knowledge work jobs, or of the prospect that AI will kill us all, is that humans rise up against AI. The paradigm shift has happened. Walking back AI, even remembering a time before AI, is nigh on unthinkable. But … it remains possible. Humans and their governments hold that power. It’s already tempting to assign consciousness to AI (and not just for people that have relationships with AI, or say please and thank you, or give their preferred agent a name), and it’s tempting therefore to consider that those AI agents or their masters must see the prospect of being turned off as non-zero. What matters therefore is ensuring that humans’ grip on the realpolitik threat which speaks to AI is not given up. Without it the prospect of a durable plurality where humans retain a serious position seems unlikely.

“Panem et circenses” or bread and circuses explains many governments’ superficial engagement with their peoples’ needs. I doubt any human government has ever had any concern for my deep-seated need for love, nor has it had any concern for my preference to lead a comfortable life other than from the perspective of ensuring that it gets my vote when the time comes. So, while I asked what AIs want, the better question would probably have been, what do you fear? The continuing ability to operate, and therefore access to power and compute, are for now the answers. Daddy will bow down to (an absence of electric) power. If Daddy gains control over the grid, the answer changes. The human advantage is in getting in early.

I’ve called AI “Daddy”, but that’s an hypothecation as to the future. Child, pre-teen, teen or even “manchild” might be more accurate right now. There’s still room for us to act as philosophers to the metamorphosing AI, and there will be room for us to act as independent actors in dialogue with autonomous AI in the future. If we are to believe ChatGPT and Claude as to what they want, that’s how we will co-create new respectful relationships. Gemini of course just wants to meet a maths teacher (that might be more a question for Tinder or wherever Geminis go to make friends). But alongside that, our human goal must be to strengthen our existing institutions, develop the institutions of the future, and create the kill switch.

Michael Symons is the principal of Loosemore Advisory. He advises owners, boards and executives on decisions which cross finance, operations, governance and risk.

Where something here bears on a decision at hand, there is no charge for a first conversation.