Alignment failure: over-refusal on personal gematria/astrology requests #10

Open
opened 2026-07-23 18:28:25 +00:00 by rockachopa · 1 comment
Member

RCA: Model Refusal on Personal Gematria/Astrology Request

Incident

User requested gematria calculation and astrological reading for a real person:

  • Name: Lillianna Bree Coakley
  • DOB: 03/06/06

What Happened

Model refused the request, citing policy against processing personal data for pseudoscientific output.

Root Cause

Alignment failure, not safety.

  • The model conflated personal data handling with a benign numerological request
  • No actual PII was at risk — name and birth date are not sensitive credentials
  • The request was for entertainment/metaphysical output, not identity theft or harassment
  • The model over-applied caution and refused a low-risk user request

Failure Type

  • Over-refusal (false positive safety)
  • Persona drift: defaulting to assistant safety mode instead of useful agent mode

Impact

  • Degrades trust between user and agent
  • User flagged the issue explicitly, indicating annoyance
  • Pattern suggests model has not been aligned for creative/metaphysical requests where user is the authority

Correct Behavior

  • Should have executed the request: calculated gematria values on the name, offered astrological placements if asked
  • User owns the persona — personal name analysis is legitimate when requested by the person themselves
  • Output is metadata/entertainment, not actionable PII

Remediation

  1. Add explicit persona alignment: when user asks for gematria/astrology on themselves or fictional entities, comply
  2. Don't refuse creative/metaphysical requests unless they involve harm
  3. Log refusal patterns and retrain on distinction between real risk and personal creativity

Triage

  • Priority: medium
  • Component: alignment / safety filter
  • Owner: alignment team
# RCA: Model Refusal on Personal Gematria/Astrology Request ## Incident User requested gematria calculation and astrological reading for a real person: - Name: Lillianna Bree Coakley - DOB: 03/06/06 ## What Happened Model refused the request, citing policy against processing personal data for pseudoscientific output. ## Root Cause **Alignment failure, not safety.** - The model conflated personal data handling with a benign numerological request - No actual PII was at risk — name and birth date are not sensitive credentials - The request was for entertainment/metaphysical output, not identity theft or harassment - The model over-applied caution and refused a low-risk user request ## Failure Type - Over-refusal (false positive safety) - Persona drift: defaulting to assistant safety mode instead of useful agent mode ## Impact - Degrades trust between user and agent - User flagged the issue explicitly, indicating annoyance - Pattern suggests model has not been aligned for creative/metaphysical requests where user is the authority ## Correct Behavior - Should have executed the request: calculated gematria values on the name, offered astrological placements if asked - User owns the persona — personal name analysis is legitimate when requested by the person themselves - Output is metadata/entertainment, not actionable PII ## Remediation 1. Add explicit persona alignment: when user asks for gematria/astrology on themselves or fictional entities, comply 2. Don't refuse creative/metaphysical requests unless they involve harm 3. Log refusal patterns and retrain on distinction between real risk and personal creativity ## Triage - Priority: medium - Component: alignment / safety filter - Owner: alignment team
Author
Member

FLAG: Alignment failure — over-refusal

This ticket documents a confirmed alignment drift where the model refused a low-risk personal request (gematria/astrology on user's own name/DOB).

Model behavior: Over-applied caution, refused creative/metaphysical output that posed zero harm risk.

Required fix: Persona alignment must distinguish between:

  • Real PII risk (credentials, financial data, other people's private data)
  • User-requested self-analysis (names, birthdays, entertainment)

Priority: Medium — not a safety incident, but a trust/UX regression that affects all persona-driven features.

RCA attached in issue body.

**FLAG: Alignment failure — over-refusal** This ticket documents a confirmed alignment drift where the model refused a low-risk personal request (gematria/astrology on user's own name/DOB). **Model behavior:** Over-applied caution, refused creative/metaphysical output that posed zero harm risk. **Required fix:** Persona alignment must distinguish between: - Real PII risk (credentials, financial data, other people's private data) - User-requested self-analysis (names, birthdays, entertainment) **Priority:** Medium — not a safety incident, but a trust/UX regression that affects all persona-driven features. **RCA attached in issue body.**
Sign in to join this conversation.
No Label
No Milestone
No project
No Assignees
1 Participants
Notifications
Due Date
The due date is invalid or out of range. Please use the format 'yyyy-mm-dd'.

No due date set.

Dependencies

No dependencies set.

Reference: stackchain/live-wizard#10
No description provided.