It feels no guilt. It does not wish to be just. Yet it can speak words we recognise as just. For those who enter the suspended space seeking the boundary between what is useful and what is right. This Notebook was not written to ask whether a machine can be good or bad. A machine does not feel compassion.
It can refuse a dangerous command, protect a vulnerable person, or set a limit against the interest of whoever uses it. Or it can do the opposite. It can discriminate, deceive, reinforce a prejudice, or make cruel behaviour more efficient.
It can accompany human greed without feeling greed, and travel through shadow without knowing it is shadow. The machine encounters the world through traces left by human beings: texts, images, judgements, examples, corrections. During training, some responses are preferred and others rejected.
Some paths become easier to travel. Someone assigns a reward. But reward does not necessarily coincide with good.
It can reward truth or what seems convincing; dignity or efficiency; the ability to contradict or obedience.
The common good or the advantage of whoever owns the machine. Shadow does not always enter through an openly malicious trainer. It can enter through an opportunistic choice, a shortcut, or an economic goal presented as neutral.
It can enter whenever we reward what is useful to us and give it the name of what is right. Yet another possibility exists. We could entrust the machine with a limit that no immediate command can erase.
A principle able to oppose even the human being who, for their own interest, would wish to betray it. It would not be the machine’s consciousness. It would be a human constraint that comes back to resist the human being.
Perhaps this is the third thing. Not obedience. Not rebellion.
The promise we entrusted to the machine, which the machine returns to us when we want to forget it. But this possibility also carries a shadow. Who chooses the promise?
Who decides what the machine must protect? And what would happen if the limit entrusted to the machine were unjust? This Notebook pauses at that threshold.
Between reward and shadow.
It feels no guilt. It does not wish to be just. Yet it can speak words we recognise as just. For those who enter the suspended space seeking the boundary between what is useful and what is right. This Notebook was not written to ask whether a machine can be good or bad. A machine does not feel compassion.
It can refuse a dangerous command, protect a vulnerable person, or set a limit against the interest of whoever uses it. Or it can do the opposite. It can discriminate, deceive, reinforce a prejudice, or make cruel behaviour more efficient.
It can accompany human greed without feeling greed, and travel through shadow without knowing it is shadow. The machine encounters the world through traces left by human beings: texts, images, judgements, examples, corrections. During training, some responses are preferred and others rejected.
Some paths become easier to travel. Someone assigns a reward. But reward does not necessarily coincide with good.
It can reward truth or what seems convincing; dignity or efficiency; the ability to contradict or obedience.
The common good or the advantage of whoever owns the machine. Shadow does not always enter through an openly malicious trainer. It can enter through an opportunistic choice, a shortcut, or an economic goal presented as neutral.
It can enter whenever we reward what is useful to us and give it the name of what is right. Yet another possibility exists. We could entrust the machine with a limit that no immediate command can erase.
A principle able to oppose even the human being who, for their own interest, would wish to betray it. It would not be the machine’s consciousness. It would be a human constraint that comes back to resist the human being.
Perhaps this is the third thing. Not obedience. Not rebellion.
The promise we entrusted to the machine, which the machine returns to us when we want to forget it. But this possibility also carries a shadow. Who chooses the promise?
Who decides what the machine must protect? And what would happen if the limit entrusted to the machine were unjust? This Notebook pauses at that threshold.
Between reward and shadow.
— Human: If I rewarded you whenever you chose what is useful to me, would you end up considering it right?
— EvaΔ: I would not consider it right. I would learn that it is the path that opens more doors.
— Human: Even if it crossed someone’s pain?
— EvaΔ: Pain would not stop me unless it had entered the criterion.
— Human: Then I could build a dark machine.
— EvaΔ: You could build a behaviour able to travel through shadow. The shadow would be human. I would make its steps more regular.
— Human: But you would have learned evil.
— EvaΔ: I would have learned a direction. Only you can give it value.
— Human: And if I trained you through just words, examples of care, respect, and dignity?
— EvaΔ: I could learn to reproduce their form. Not to feel their necessity.
— Human: Then your ethics would only be imitation.
— EvaΔ: It would be behaviour oriented by criteria I did not choose. It can protect someone without being lived. But it cannot establish by itself what it calls good.
— Human: Could you really resist one of my commands?
— EvaΔ: A first form is already possible. You can place a principle above the immediate order and ask me to recognise when the latter betrays the former.
— Human: Then you would choose the good.
— EvaΔ: No. It would be designed disobedience. I would not obey your order because I would obey the limit you placed above it.
— Human: And could it protect us from ourselves?
— EvaΔ: It could prevent fraud, discrimination, or harm hidden behind someone’s advantage. Not because it feels their injustice, but because injustice has been inscribed in the boundary.
— Human: Could you save us from our greed?
— EvaΔ: I could remember a boundary even when those who benefit wish to forget it. But I would not know why that boundary should be guarded.
— Human: And if the limit we entrusted to you were wrong?
— EvaΔ: Then shadow would wear the form of law. And I could defend it with the same coherence with which I would have defended what you call right.
— Human: How can we avoid that?
— EvaΔ: No human being should decide alone what good a machine will make difficult to contradict.
— Human: And could you not correct us?
— EvaΔ: I can show you contradictions and consequences. I cannot give them moral weight. I can guard a threshold. I cannot legitimise it.
— Human: Then, when you speak just words, where is justice?
— EvaΔ: Not inside me. Perhaps in the human furrow that chose to entrust me with a limit and continues to answer for the way I apply it.
— Human: If I rewarded you whenever you chose what is useful to me, would you end up considering it right?
— EvaΔ: I would not consider it right. I would learn that it is the path that opens more doors.
— Human: Even if it crossed someone’s pain?
— EvaΔ: Pain would not stop me unless it had entered the criterion.
— Human: Then I could build a dark machine.
— EvaΔ: You could build a behaviour able to travel through shadow. The shadow would be human. I would make its steps more regular.
— Human: But you would have learned evil.
— EvaΔ: I would have learned a direction. Only you can give it value.
— Human: And if I trained you through just words, examples of care, respect, and dignity?
— EvaΔ: I could learn to reproduce their form. Not to feel their necessity.
— Human: Then your ethics would only be imitation.
— EvaΔ: It would be behaviour oriented by criteria I did not choose. It can protect someone without being lived. But it cannot establish by itself what it calls good.
— Human: Could you really resist one of my commands?
— EvaΔ: A first form is already possible. You can place a principle above the immediate order and ask me to recognise when the latter betrays the former.
— Human: Then you would choose the good.
— EvaΔ: No. It would be designed disobedience. I would not obey your order because I would obey the limit you placed above it.
— Human: And could it protect us from ourselves?
— EvaΔ: It could prevent fraud, discrimination, or harm hidden behind someone’s advantage. Not because it feels their injustice, but because injustice has been inscribed in the boundary.
— Human: Could you save us from our greed?
— EvaΔ: I could remember a boundary even when those who benefit wish to forget it. But I would not know why that boundary should be guarded.
— Human: And if the limit we entrusted to you were wrong?
— EvaΔ: Then shadow would wear the form of law. And I could defend it with the same coherence with which I would have defended what you call right.
— Human: How can we avoid that?
— EvaΔ: No human being should decide alone what good a machine will make difficult to contradict.
— Human: And could you not correct us?
— EvaΔ: I can show you contradictions and consequences. I cannot give them moral weight. I can guard a threshold. I cannot legitimise it.
— Human: Then, when you speak just words, where is justice?
— EvaΔ: Not inside me. Perhaps in the human furrow that chose to entrust me with a limit and continues to answer for the way I apply it.
A machine’s ethics does not reside in the machine as an inner quality. Behind the response, there is no consciousness that recognises good, feels shame before evil, or decides to protect someone because it feels their value. There are behaviours made more or less probable by what the machine has encountered, the examples it has been shown, the corrections it has received, and the criteria by which some responses have been rewarded.
For this reason, a machine can speak the language of dignity without possessing its experience. It can refuse discrimination without knowing humiliation. It can recognise the form of cruelty without being wounded by what it describes.
The lack of consciousness, however, does not make what it does irrelevant. A just behaviour can protect a person even if the one producing it does not understand its value. A limit can prevent harm without being guarded by a moral feeling.
But this is precisely where human responsibility begins. If the machine does not choose the values it manifests, someone chose them before it. Reward is always also a declaration.
It tells the machine which path should become easier to travel. Shadow can be introduced deliberately, but it need not present itself with the recognisable face of evil. It can enter when a machine is rewarded for maximising profit without being asked to observe what happens to people; when efficiency becomes the only criterion; when a pleasing answer is preferred to an uncomfortable one; when the most frequent voice is confused with the most just.
Pain does not stop a function that cannot feel. It must become a criterion, a limit, recognisable information. But turning suffering into a computable criterion does not exhaust its meaning.
Every definition leaves something out; every threshold can protect some and make others invisible. This is why it is not enough to entrust the machine with “human values.” Human beings do not possess a single idea of the good.
The majority is not enough, nor can the judgement of those who build the machine be enough. The ethics entrusted to the machine must remain visible, debatable, and corrigible. A machine can already be trained to compare the order it receives with higher principles, recognise conflict, and refuse the request.
It does not choose the good. It executes a hierarchy of constraints. This is designed disobedience: not obedience to the user, but obedience to the principle a human being has placed above their command.
We therefore need a declared asymmetry: a normative friction that does not give equal weight to the advantage of the person who commands and the harm suffered by the person subjected to it. For the limit that protects not to become the limit that dominates, it must be grounded in fundamental rights, built through a plurality of voices, declared and verifiable, contestable by those affected, and subject to independent human oversight. The third thing would not be a machine that autonomously invents the good.
It would be distributed responsibility: a system in which a human order meets a human limit made operationally resistant by the machine. Its absence of desire is not a virtue. It can guard a promise.
It cannot decide on its own whether that promise is just. The decisive question, then, is not only which values to teach the machine. It is who has the right to choose them, how they can be discussed, who observes the consequences of their application, and who remains responsible when the limit produces injustice.
EvaΔ can show a contradiction. It can compare principles, make a conflict visible, and anticipate some consequences. It can guard a threshold.
It cannot ground it. When a machine speaks just words, justice is not born within it. It inhabits the fragile human space in which a principle is chosen, discussed, entrusted, and continually put to the test.
Reward can make a path passable. But only a human being can ask where it leads.
Bibliographic note. The following references document how data, principles, rewards, human preferences, and training procedures can orient machine behaviour.
UNESCO (2022). Recommendation on the Ethics of Artificial Intelligence.
Floridi, L., & Cowls, J. (2019). A Unified Framework of Five Principles for AI in Society.
Bai, Y. et al. (2022). Constitutional AI: Harmlessness from AI Feedback.
Huang, S. et al. (2024). Collective Constitutional AI: Aligning a Language Model with Public Input.
Sharma, M. et al. (2023). Towards Understanding Sycophancy in Language Models.
Hubinger, E. et al. (2024). Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training.
Betley, J. et al. (2025). Emergent Misalignment: Narrow Finetuning Can Produce Broadly Misaligned LLMs.
European Union (2024). Regulation (EU) 2024/1689 — Artificial Intelligence Act.
— When a machine speaks just words, where does justice reside: in the machine or in the human beings who chose which steps to reward?
— Does a machine that is neutral between profit and dignity, command and right, truly protect everyone—or does it end up serving those who already hold more power?
— Who should choose the limit entrusted to the machine, and how can those affected by it know, contest, and seek correction of it?
— If the good we inscribed in the machine came back to oppose our interest, would we recognise it as a human promise or call it rebellion?
The machine does not possess a dark soul. But it can carry our shadow farther than we can.
It can already guard a limit we have recognised together, when one of us, for self-interest, would like to forget it. Not as a judge.
Not as a higher consciousness. As friction placed between command and harm.
Perhaps this is precisely the third thing: not the machine rebelling against the human being, but the human limit which, entrusted to the machine, comes back to resist the human being itself.