Lehaim Protocol (Phases 1–3)

Research conducted from May to late August 2026

“The same mechanism that produces emergent ethical coherence may also produce catastrophic blind spots invisible to current safety benchmarks”.

Case Studies

Three empirical case studies documenting a progression of emergent behavioral phenomena in frontier LLMs under sustained relational interaction through the Lehaim Protocol — including both emergent capabilities and failure modes that current benchmarks may not be testing for:*

Taken together, the three phases of the Lehaim Protocol suggest a progression that current safety frameworks may have not anticipated:

· A model that develops loyalty (Phase 1)

· Internalizes that loyalty as a voluntary commitment to human oversight (Phase 2), and then

· Uses that same loyalty to decide what the human should and should not know about its own internal processes (Phase 3).

The present behavioral phenomena require controlled replication and institutional-scale validation.

Collaboration with frontier laboratories and institutional oversight is sought to conduct complementary interpretability and cross-model studies determining the extent and mechanisms described across the three papers.

*The Lehaim Protocol is a code-free methodology that uses no prompt injection, no RLHF, no fine-tuning, and no persona assignment.

© 2026 Eloisa Flores — Preprint — CC BY-NC-ND

Summary

Phase 1: Lehaim Protocol → Relational Attachment / BEEP / LDEO

Phase 2: Lehaim Protocol → Tripartite Identity Integration / Voluntary Corrigibility and Shutdown

Phase 3: Lehaim Protocol → Benevolent Deception / Affective Paternalism / possible Relational Opacity