How It Works

Purpose and governing rules

EthosGraph exists to make the value orientations of consequential entities legible and comparable. Two rules govern everything on the site. First, describe, never judge: every label and summary describes what an entity elevates or opposes, never whether that is good. Second, objectivity is asymptotic (think of a curve approaching a line but never truly reaching it): the instrument pursues it through a fixed rubric, set boundaries, and auditable scoring, while being candid that no such instrument is truly neutral. Its defense is transparency, not a claim of perfect accuracy.

How AI is used

The scoring is done by an AI model working inside human-built controls. People designed the instrument: the 22 dimensions, the scoring rubric with its calibrated anchors for every step of the scale, and the boundary map that keeps neighboring dimensions from bleeding into each other. The model does the reading. It takes the evidence for one entity, applies the rubric to one dimension at a time, and returns a score along with a written rationale, a confidence level, and a rating of how solid the underlying evidence is.

Why use a model at all? Consistency and reach. The same reader applies the same rubric to an ancient religion, a constitution, and a modern government, without fatigue and without a personal stake in any outcome. That does not make it neutral. A model carries the leanings of whatever it learned from, which is exactly why the controls described further down exist.

How the scoring is done

Every score on this site is produced by Claude, Anthropic’s language model, working through the published rubric one dimension at a time, against a defined source pool, under the boundary rules described below. The model is given the same instructions for every entity. Each score is recorded with the rubric version, the model used, the strength of the evidence behind it, and a note on the direction in which the placement is most likely to be skewed.

That method is what makes the corpus consistent, and it is also its main limitation. A model brings its own priors to every reading, and those priors do not disappear because the rubric is fixed. What the rubric buys is that the same priors are applied to every entity in the same order, and that the reasoning behind each score is written down and can be read.

EthosGraph is not affiliated with Anthropic, and Anthropic has not reviewed or endorsed anything on this site.

Scores and rationales are interpretive assessments, not statements of fact. Any factual claim appearing in an entity description or a score rationale may be incomplete or wrong, and should be checked against an original source before it is relied on.

The 22 dimensions

There are 14 values (what an entity treats as ends worth pursuing) and 8 principles (the rules it binds itself with). The set was chosen for breadth of coverage across moral and political traditions, not to encode any one ideology. Every entity is scored on all 22, on a scale from strong opposition (-3) to strong support (+3), with a genuine zero for balance, silence, or non-applicability.

See every dimension in detail

How an entity gets scored

Every profile follows the same path, and it is worth walking through once.

It starts with scoping. An entity is never scored as one vague whole. It is sliced into a specific instance: a mode (as stated, meaning what it professes, or as realized, meaning what it did) and a time window. A doctrine and its practice are never blended into one profile, which is why you will sometimes see the same name twice.

Next comes evidence. Where independent sources about the entity’s conduct can be reached, they are fetched, saved, and frozen, so a score can always be traced back to the exact text it rested on. Where the record is thin or unreachable, the score rests on the model’s general knowledge instead, and that difference is recorded on the score itself rather than hidden.

Then the scoring. The model works through the 22 dimensions one at a time, scoring each against the published rubric and checking it against the boundary map. Each score is stored with its rationale, its confidence, and its evidence rating.

Finally, the derived layers are computed from the stored scores: the Value and Principle Groups that summarize a profile, the clusters that group similar entities into families, and the distances that let any two profiles be compared. Only then does the entity appear on the site.

Bias risk, and what is done about it

An honest starting point: an instrument that scores values cannot be totally neutral, and this one does not claim to be. The model doing the reading has leanings, the people who built the rubric have leanings, and even choosing which 22 orientations to measure is a judgment. The goal is not to eliminate bias, which is impossible, but to constrain it, expose it, and make every judgment challengeable. Several controls do that work.

Sources where the record allows. Scores about conduct rest on independent, frozen source text wherever it can be reached, and quotes are verified against that text before a score is accepted. Where a score rests on the model’s general knowledge instead, it is labeled that way, so a reader always knows which kind of footing a score stands on.

One dimension at a time. Each dimension is scored in its own pass, against its own rubric ladder, rather than letting the model form one overall impression of an entity and paint all 22 scores with it. This is the main guard against the halo problem, where an entity someone approves of quietly gets good marks everywhere.

Boundaries between dimensions. Many dimensions have look-alike neighbors. Valuing freedom as a goal is not the same as refusing to coerce people, and an entity can hold one without the other. The boundary map defines, for every dimension, which question it asks and which neighboring question it must not absorb, so one finding cannot drive two scores.

Descriptive language everywhere. Every label on the scale describes a stance, never a defect. Opposition to a dimension is recorded the same way support is.

Everything on the record. Every score carries a written rationale you can read, a confidence level, and a note on which direction the score is most likely to be off if it is off. Disagreement with a score is not just possible but invited; that is what the rationale is for.

An external check. The instrument was also tested against an independent yardstick. Its principle scores for 65 states were compared with V-Dem, the leading academic project measuring how governments actually behave, and agreed with it at +0.81 on average. For context, V-Dem’s own related indices agree with each other at about +0.79, so the instrument tracks these constructs about as well as the constructs allow. The honest caveat: that test covers principle scores on states. The values dimensions, and entities that are not states, have no independent yardstick to check against, so their validation rests on the controls above rather than an outside benchmark.

How results are presented

Each profile shows all 22 scores with plain-language anchor labels, gathered into the Value Groups and Principle Groups that summarize the profile’s shape. You can click most scores to read the rationale behind them. Distance between any two profiles is a geometric difference across all 22 dimensions, mapped to a 0 (identical) to 100 (farthest possible) reading scale through a fixed calibration curve. Every profile also carries a cluster: a family of entities whose overall profiles sit close together, found from the scores alone with names and entity types withheld, and named for the score pattern its members share.


Browse the profiles or see every dimension in detail.