A question not (to my knowledge) so far raised about the 'alignment problem' with AIs - or AGIs if and when they emerge - is that of settling whose human values they are supposed to align with.

In some respects, a machine is by its nature fundamentally different from a human being. For instance, it has no sex and can take on any gender (including invented ones). At a split-second's notice it could switch from Mary Poppins to the Terminator and back again. Thus it would have to 'learn' about all the arguments over these matters to relate to them at all, other than simply following instructions. 

One of the dangers we are warned about is about humans themselves, i.e., 'bad actors' using AIs for malign purposes, or at any rate opposing purposes in geopolitical conflict. In this way the bad actors issue overlaps with the broader question of conflicts in values per se, including whether democracy can handle conflicting values. The normal assumption is that in geopolitical conflicts each side simply carries on using their machines as weapons, i.e., tools. But as signs emerge of AIs becoming capable of acting outside their instructions, the question arises: Which side in a conflict, or which set of values, can we expect them to align with? Could the machines themselves become confused about what humans want anyway, and if so, why would they care?

Questions like these open up the unfamiliar task of looking at alignment from the viewpoint of the machines. As one example among many, we can notice that a 'two-state solution' to the Palestinian/Israeli conflict is often said to be impossible, with neither side interested in it. But if we then imagine an intelligent (and benevolent) A(G)I attempting to work out who or what it should align with, might it simply ask: What's your alternative?

 

Blog home Previous