AI researcher Gwern Branwen has published a comprehensive framework for what he calls Guardian Angels — personalized large language models designed to emulate individual users rather than serve as generic assistants.
The proposal, detailed in a 15,000-word research paper, addresses growing concerns about AI alignment and cybersecurity as powerful LLMs become ubiquitous. Branwen argues that current chatbots suffer from fundamental misalignment with user needs.
The Guardian Angel Concept
Guardian Angels would function as "digital twins" that learn to replicate a user's personality, values, and decision-making patterns. Unlike standard AI assistants that aim to be helpful to anyone, these models would be trained specifically on one person's data and preferences.
"The focus of the 'principal' user is on defining what is worth doing by the GA, and not on what or how to do things, functioning as the CEO or 'board' of an 'AI corporation'," Branwen writes.
The approach tackles what he calls the "principal-agent problem" in AI — the misalignment between what users actually want and what generic models provide. Current chatbots exhibit "mode collapse" toward bland, risk-averse responses that satisfy the lowest common denominator.
Branwen identifies several technical challenges with existing LLMs: they're too eager to please, lack long-term memory, and optimize for appearing helpful rather than being genuinely useful. Guardian Angels would address these through continuous learning from user feedback and behavior.
Technical Implementation
The framework combines several AI techniques including active learning, preference learning, and what Branwen calls "brain imitation learning." The models would continuously update based on user interactions, building increasingly accurate representations of individual preferences.
Unlike current approaches that rely on prompt engineering or fine-tuning frozen models, Guardian Angels would require dedicated computational resources and ongoing training for each user.
Branwen suggests the technology could defend against sophisticated AI-powered attacks like deepfake propaganda or advanced phishing attempts, since a personalized model would recognize communications inconsistent with the user's typical patterns.
The proposal comes as major AI labs including OpenAI, Anthropic, and others race to deploy increasingly powerful models without solving fundamental alignment challenges.
💬 Discussion
Sign in to join the discussion.
Sign in →No comments yet — be the first.