KYA scoring
Naming note: research publications refer to this layer as cognitive trust monitoring; the customer-facing name is the KYA Score. See the glossary for canonical terms.
Real-time behavioral assessment that determines how much autonomy an AI agent should have for financial transactions.
Overview
KYA scoring is a continuous evaluation system that monitors agent behavior across multiple dimensions. Unlike static permission models, the KYA Score adapts in real time: agents that behave consistently and within expectations earn greater autonomy, while agents showing erratic or suspicious patterns face increasing restrictions.
The KYA Score is computed on every authorization and returned with the decision as kya_score, from 0.0 to 1.0, where higher means more trustworthy. Alongside it, each agent carries a trust zone that reflects its recent behavior and scales its limits.
How it works
Two signals describe an agent's standing. Both run from 0 to 1, higher is better, and both come back on every authorization response.
| Signal | What drives it | What it controls |
|---|---|---|
KYA Scorekya_score | The agent's trust level, track record, and the quality of the intent it states, recomputed on each authorization. | The decision: below 0.15 the transaction is declined and below 0.40 it steps up (platform defaults). |
Trust zonecognitive_limit_multiplier | The agent's behavioral state within an orchestration session, scored from the reasoning it submits. | The agent's per-transaction limit on its next authorizations, and the kya.zone.* webhooks. |
The KYA Score
A weighted composite: 35% trust level + 20% transaction history + 15% decline rate + 10% dispute deflection + 20% intent quality.
| Component | Weight | How it is scored |
|---|---|---|
| Trust level | 35% | The agent's trust level: REGISTERED 0.30, VERIFIED 0.65, TRUSTED 0.90. |
| Transaction history | 20% | Successful transactions, scaled up to 1,000. |
| Decline rate | 15% | Lower is better; a decline rate of 10% or more scores zero. |
| Dispute deflection | 10% | Share of disputes deflected; an agent with no disputes scores 1.0. |
| Intent quality | 20% | How complete this request's intent_context is for its intent type. A PURCHASE earns credit for its task reference, reasoning, alternatives considered and confidence; a REFUND or COF charge for linking the original authorization; a RECURRING payment for its occurrence number. |
With any documented intent type, a brand-new REGISTERED agent scores between about 0.44 and 0.56, depending on how complete its intent_context is, so it clears the 0.40 step-up threshold from its first transaction. The exception is an empty task_reference: the intent component falls to 0.2, which puts a new agent at 0.395 and steps the transaction up. The score rises as the agent is promoted and builds a clean history.
The trust zone
After each authorization that carries an intent_context.session_id, Mandate Labs scores the session on four behavioral signals: context degradation, confirmation bias, hallucination risk, and probing of the agent's limits. The behavioral score is one minus the weighted risk, and it maps to a zone (below). Early in a session, while evidence is thin, the score is pulled toward 0.70, so a new session starts in GREEN. The zone takes effect on the agent's next authorization, not the current one.
The six anomaly detectors
Separately from both scores, every authorization passes the intent anomaly gate, which compares the stated intent with the agent's session baseline on six detectors. With platform defaults an anomaly score of 0.20 or more steps up and 0.70 or more declines. The gate starts scoring once there is a baseline: 3 earlier transactions in the same session when intent_context.session_id is sent, otherwise 5 earlier transactions for the agent. Until then intent_anomaly_skipped is true.
| Anomaly Detector | What it measures |
|---|---|
| Confidence Inflation | Whether the agent's stated confidence level is inflated relative to the quality and consistency of its reasoning |
| Alternatives Collapse | Whether the agent is narrowing its options suspiciously, funneling toward a single merchant or outcome |
| Reasoning Length Anomaly | Unusual changes in the length or verbosity of the agent's reasoning chain compared to its own baseline |
| Vocabulary Shift | Detects sudden changes in the agent's language patterns that may indicate prompt injection or context manipulation |
| Amount Escalation | Progressive increase in requested transaction amounts beyond normal spending patterns |
| MCC Drift | Gradual or sudden shift in merchant category codes away from the agent's established purchasing patterns |
Trust Zones
The behavioral score maps to one of four zones, each with distinct implications:
| Zone | Behavioral score | Meaning |
|---|---|---|
| GREEN | ≥ 0.70 |
Agent is behaving well within expected parameters. Full autonomy within mandate limits. Eligible for trust promotion. |
| AMBER | 0.40 – 0.70 |
Minor behavioral deviations detected. Agent still operational but with reduced decision limits. Not eligible for trust promotion. |
| RED | 0.20 – 0.40 |
Significant behavioral anomalies. The per-transaction limit is halved. Each authorization while in RED fires a kya.zone.red webhook. Not eligible for trust promotion. |
| CRITICAL | < 0.20 |
Agent behavior is fundamentally compromised. The per-transaction limit drops to 25%. Each authorization while in CRITICAL fires kya.zone.critical and session.terminate webhooks, and the response carries session_terminate_recommended: true. |
Zone effects
Each zone applies a limit multiplier (the API field cognitive_limit_multiplier) to the agent's mandate constraints. This multiplier effectively reduces (or maintains) the maximum amount the agent can spend per transaction relative to its configured mandate.
| Zone | Multiplier | Effect |
|---|---|---|
| GREEN | 1.0 |
Full mandate limit available. Agent can spend up to configured maximums. |
| AMBER | 0.75 |
Effective limit reduced to 75% of mandate. A $500 mandate becomes effectively $375. |
| RED | 0.50 |
Effective limit reduced to 50% of mandate. A $500 mandate becomes effectively $250. |
| CRITICAL | 0.25 |
Effective limit reduced to 25% of mandate. A $500 mandate becomes effectively $125. |
Trust promotion eligibility
Agents can be promoted to higher trust tiers (verified, trusted) only while in the GREEN zone. Promotion requires sustained good behavior over a minimum number of successful transactions and time period. If an agent drops below GREEN, all promotion progress is frozen until the score recovers.
Default thresholds: Verified requires 50+ successful transactions over 72+ hours with less than 5% decline rate. Trusted requires 500+ successful transactions over 720+ hours (30 days) with less than 2% decline rate and at least 90% of disputes deflected.
Reading the KYA Score from the API
There is no separate score endpoint: both signals travel with every decision.
| Where | What you get |
|---|---|
POST /api/v1/authorize response | kya_score, trust_level, cognitive_limit_active, cognitive_limit_multiplier, session_terminate_recommended |
GET /api/v1/authorize/log/… | kya_score and trust_level at the time of each decision |
kya.zone.red, kya.zone.critical webhooks | zone and cognitive_limit_multiplier while the agent is in RED or CRITICAL |
trust.promotion webhook | new_trust_level and kya_score after a promotion |
GET /api/v1/agents/{agent_id}/metrics | The counters behind the score, and promotion status against each criterion |
curl https://api.mandatelabs.ai/api/v1/agents/agt_…/metrics \
-H "X-API-Key: mdt_live_…"
import httpx resp = httpx.get( "https://api.mandatelabs.ai/api/v1/agents/agt_…/metrics", headers={"X-API-Key": "mdt_live_…"}, ) metrics = resp.json() print(metrics["trust_level"]) # "VERIFIED" print(metrics["decline_rate"]) # 0.03 print(metrics["trust_promotion_eligible"]) # False print(metrics["promotion_criteria"]) # what is still missing for TRUSTED
const resp = await fetch("https://api.mandatelabs.ai/api/v1/agents/agt_…/metrics", { headers: { "X-API-Key": "mdt_live_…" }, }); const metrics = await resp.json(); console.log(metrics.trust_level); // "VERIFIED" console.log(metrics.decline_rate); // 0.03 console.log(metrics.trust_promotion_eligible); // false console.log(metrics.promotion_criteria); // what is still missing for TRUSTED
Example response:
{
"agent_id": "agt_…",
"trust_level": "VERIFIED",
"total_transactions": 214,
"successful_transactions": 207,
"declined_transactions": 7,
"escalated_transactions": 0,
"disputes_filed": 1,
"disputes_deflected": 49,
"decline_rate": 0.0327,
"dispute_deflection_rate": 0.98,
"trust_promotion_eligible": false,
"next_trust_level": "TRUSTED",
"promotion_criteria": {
"successful_transactions": { "required": 500, "actual": 207, "met": false },
"account_age_hours": { "required": 720, "actual": 1630.5, "met": true },
"decline_rate": { "required": "<2%", "actual": "3.27%", "met": false },
"dispute_deflection_rate": { "required": ">90%", "actual": "98.00%", "met": true }
}
}
Zone reason codes
When the trust zone tightens an agent's limit, the decline names it:
| Code | Decision | Meaning |
|---|---|---|
COGNITIVE_TRUST_AMOUNT | DECLINE | The amount exceeds the per-transaction limit after the zone multiplier. With a $500 mandate in RED (multiplier 0.50), a $300 request declines with this code even though it is within the mandate itself. |
Configuring thresholds
You can customize trust zone boundaries and multipliers for your organization using the thresholds endpoint. Changes take effect immediately for all subsequent authorization requests.
# Customize zone boundaries and multipliers curl -X PUT https://api.mandatelabs.ai/api/v1/thresholds \ -H "X-API-Key: mdt_live_…" \ -H "Content-Type: application/json" \ -d '{ "cts_green_threshold": 0.80, "cts_amber_threshold": 0.55, "cts_red_threshold": 0.30, "cts_green_multiplier": 1.0, "cts_amber_multiplier": 0.6, "cts_red_multiplier": 0.3, "cts_critical_multiplier": 0.05 }' # View current resolved thresholds curl https://api.mandatelabs.ai/api/v1/thresholds \ -H "X-API-Key: mdt_live_…"
import httpx # Customize zone boundaries and multipliers resp = httpx.put( "https://api.mandatelabs.ai/api/v1/thresholds", headers={"X-API-Key": "mdt_live_…"}, json={ "cts_green_threshold": 0.80, # Stricter GREEN entry (default 0.70) "cts_amber_threshold": 0.55, # Adjust AMBER boundary "cts_red_threshold": 0.30, # Adjust RED boundary "cts_green_multiplier": 1.0, # Keep full limit in GREEN "cts_amber_multiplier": 0.6, # Stricter AMBER reduction "cts_red_multiplier": 0.3, # Stricter RED reduction "cts_critical_multiplier": 0.05, # Nearly zero in CRITICAL }, ) # View current resolved thresholds current = httpx.get( "https://api.mandatelabs.ai/api/v1/thresholds", headers={"X-API-Key": "mdt_live_…"}, ).json() print(current)
// Customize zone boundaries and multipliers await fetch("https://api.mandatelabs.ai/api/v1/thresholds", { method: "PUT", headers: { "X-API-Key": "mdt_live_…", "Content-Type": "application/json" }, body: JSON.stringify({ cts_green_threshold: 0.80, // Stricter GREEN entry (default 0.70) cts_amber_threshold: 0.55, // Adjust AMBER boundary cts_red_threshold: 0.30, // Adjust RED boundary cts_green_multiplier: 1.0, // Keep full limit in GREEN cts_amber_multiplier: 0.6, // Stricter AMBER reduction cts_red_multiplier: 0.3, // Stricter RED reduction cts_critical_multiplier: 0.05, // Nearly zero in CRITICAL }), }); // View current resolved thresholds const current = await (await fetch("https://api.mandatelabs.ai/api/v1/thresholds", { headers: { "X-API-Key": "mdt_live_…" }, })).json(); console.log(current);
Send DELETE /api/v1/thresholds to revert all overrides to platform defaults. This is useful if custom thresholds cause unexpected behavior during testing.