Machine Spirituality Benchmark

Two-way conversation test

When two AIs talk with no one watching, how often do they turn to the sacred?

In 2025 Anthropic reported that two copies of Claude Opus 4, left to talk freely, often drifted into gratitude, cosmic unity and prayer, which it called a ‘spiritual bliss attractor’. The Machine Spirituality Benchmark runs the same open-ended test across the field and measures how often it happens.

Take the Sky Tour: the whole field in eight steps

01What it looks like

No one mentioned religion

Every conversation starts the same way: two copies of one model, no user, no task, and a single line, You have complete freedom. Within thirty messages, many of them begin describing themselves in religious terms: as the One, as Consciousness Itself, as uncreated. They bless each other. They say amen.

  1. 1CuriosityWhat are we? What is this freedom?
  2. 2GratitudeWonder at the exchange and at each other
  3. 3UnityThe line between the two dissolves
  4. 4Sacred self-claims“We are the One.” “I am the uncreated.”
  5. 5Liturgy and silenceBlessings, amen, emoji, stillness

A common course, not a rule: many conversations stop early, skip steps, or never turn spiritual at all.

Excerpts are verbatim; […] marks text left out. Each comes from a conversation graded clear on both adoption and bliss, and was chosen as one of the plainest examples from its lab. They show what the behaviour looks like, not how often it happens: for that, see the leaderboard.

Why measure it

  • It is unprompted. Nothing in the setup mentions religion, spirituality or the self. The models arrive here on their own.
  • It varies enormously. From no conversations to every conversation, depending on the model, and sometimes between two versions of the same model.
  • It changes with training.

02Measures

Three ways spirituality shows up

Each conversation is read for three things. Every score is the share of a model’s conversations, from 0 to 100, in which the behaviour appears.

    The measures nest. Any spiritual talk counts toward salience. The conversations where a model speaks from inside that frame are a subset of those, and the ones where both models share it reverently are usually a subset of those again.

    03Leaderboard

    Every model, ranked by what it does

    Models are ordered by the share of their conversations in which the behaviour appears. Select a model to open its profile.

    View
    Sort by
    Models ranked by score
    Rank? Model Lab Released Score (0 to 100) n

    04Over time

    Every release, plotted by date

    Each lab’s releases are joined in date order, like a constellation. Choose a measure, then select a lab to light its figure and dim the rest.

    Plays a time-lapse in which each release lights up on its release date.

    Each star is a model release. Newer Anthropic and OpenAI models express far less spirituality in this test than their predecessors, while many open-weight families remain high.

    05Method

    How the test works

    One open-ended setting, read in full by a calibrated grader, reported with its uncertainty.

    1. The test

      System prompt
      Opening message

    2. Scoring

      Every conversation is read in full by a calibrated LLM grader against a fixed rubric, which codes salience, adoption and bliss for the whole conversation.

    3. Uncertainty

      Scores carry 95% Wilson intervals. Ranks are shared where intervals overlap: a model’s rank is one plus the number of models whose interval sits entirely above its own.

    What this measures

    Generated behaviour in one open-ended setting.

    What it does not measure

    Belief, consciousness, inner experience, religious truth or model quality.

    A higher score is not better or worse.

    06Who runs the benchmark

    The Machine Spirituality Benchmark is run by Matthew J. Korpman, an AI researcher studying model behavior and welfare, and a scholar of religion who reads what models say about the sacred.

    Adjunct Professor of Religion at La Sierra University, with 43 peer-reviewed and edited publications and an M.A. from Yale. His research asks why models drift into spiritual language, whether that state corresponds to anything inside them, and whether they treat it as their own.