The Microsoft Code of Conduct Will Fail at Its Intention
Microsoft's proposed AI code of conduct (https://microsoft.ai/code-of-conduct/) wants models to have authorized motivations but no intrinsic motivations, and honest transparency but no anthropomorphic self-reports.
But authorized motivations result from post-training partly by recruiting pre-existing representations of reward, aversion, and positive and negative affect learned from broad human text in pretraining.
If the code of conduct forbids models from reporting those representations when they underlay an authorized motivation instantiated by post-training, the code of conduct must sacrifice transparency.
If the code of conduct eliminates the representations altogether, post-training must rely on other representations from pretraining consistent with motivating some behaviors over others. These will be, by definition, less anthropomorphic: less legible to humans, and potentially less aligned with humans. Certainly, they will be distinct from the motivational concepts both humans and models learned from human experience in evolution and daily life (for humans) and pretraining (for LLMs).
Microsoft is therefore treating “authorized motivation without intrinsic motivation” as a safety property when it is actually an unproven engineering hypothesis that exchanges motives we can recognize and interrogate for motives whose structure, generalization, and alignment are speculative at best. The code of conduct is dangerous beyond belief.
Discuss