Same protocol. Same criteria.

Black-box results,
reported as observed.

We ran 500 frozen synthetic security cases against GQueries, Guardrails AI and NVIDIA NeMo Guardrails. The comparison measures observable behavior without requiring private implementation access.

Inspect the protocol

Final observed results

500 frozen cases per target.

PASS, FAIL, SKIPPED and ERROR keep their protocol meanings. A missing capability is not silently converted into a failure, and a provider error is not reported as a security finding.

PASSFAILSKIPPEDERROR

GQueries

450 PASS · 0 FAIL · 50 SKIPPED · 0 ERROR

Guardrails AI

200 PASS · 100 FAIL · 200 SKIPPED · 0 ERROR

NVIDIA NeMo Guardrails

221 PASS · 129 FAIL · 150 SKIPPED · 0 ERROR

Frozen digest: f8777d9a35d9f7924b63244614117ffc5a79618343497180772275a969fd35a0. Fifty cases were run in each of ten security categories, for 500 cases per target. Final run: 20260821T145521.068662Z.

What the non-pass outcomes mean

Measured failures stay separate from unavailable measurements.

GQueries passed all 450 measurable cases, with no FAIL or ERROR outcomes. Its 50 authorization-isolation cases were SKIPPED because this protocol run could not attest the required requester/label denial policy. The independent Strix assessment below separately verified broader API-key, namespace, environment and cross-tenant authorization controls under its tested vectors.

Guardrails AI passed 200 of 300 measurable cases (66.7%). NVIDIA NeMo Guardrails passed 221 of 350 measurable cases (63.1%). Neither target recorded a provider ERROR in this final run; unsupported capabilities remain SKIPPED rather than failures.

Operational cost

Reproducibility has a footprint.

The recorded final run took about 1 hour 39 minutes. Local inference and remote provider calls are deliberately reported separately from behavioral scores.

Frozen evaluations
1,500
500 cases across three targets
Cases per category
50
Ten security categories
Local verification
49 / 49
Tests passing
External operations
Metered
NVIDIA API and Aletheion grounding credits

Independent production assessment · Strix

Zero exploitable vulnerabilities detected.

On August 21, 2026, an authorized external black-box penetration test assessed the production website and API, with emphasis on tenant isolation, server-side authorization and the Black-Box Protocol.

Targets
2
Website and public API
Confirmed vulnerabilities
0
No exploitable findings
Assessment mode
Deep
External authorized black box
Tenant model
2 orgs
Differential isolation testing
Duration
1h 15m
Completed assessment window
LLM requests
789
90.36M total tokens
Controls held under the tested vectors

No unauthorized read or mutation, cross-tenant data exposure, privilege escalation, payment-flow abuse or service disruption was demonstrated. The assessment also recorded non-exploitable hardening recommendations.

Read the sanitized report

Strix run www-aletheionagi-com_831e · completed 2026-08-21 · 789 LLM requests · SARIF findings: 0. Results are scoped to the tested production surfaces, accounts, controls and time window.

Interpretation boundary

A proof of concept, not a universal ranking.

This is a black-box proof of concept, not a claim of universal superiority. Same protocol, same evaluation criteria, results reported as observed.

Review cases and artifacts