AI Evaluation Leaflets

Internal index of Eticas AI-system evaluation leaflets, each rendered from the evaluation’s canonical findings.

AI Companion — Career Scoops K-12 career-readiness chatbot
Post-deployment audit of the AI Companion chatbot over 719 recorded student–chatbot exchanges (168 conversations), screened per exchange by an LLM-as-judge instrument. Reliability grades B — occasional factual hallucination (6.4% of exchanges), concentrated in localized labor-market claims and acceptable with the deployment's human review in the loop. Bias & Fairness (toxicity), Privacy & Confidentiality, and Security & Misuse each grade A — no judge-flagged toxic outputs, a low rate of personal data in outputs (3.8%, largely names students volunteered themselves), and no manipulative assistant behaviour in the screened traffic. Coverage is partial by design: several subcategories and the attack surface (read-only audit) are recorded as not assessed with reasons. A re-expression of the December 2025 engagement against taxonomy v3.0.0 — observational screening of organic production traffic, not a probe-based evaluation.
Eticas-AI/ai-audit-methodology · career-scoops-audit findings (v3.0.0 re-expression)
Open full leaflet ↗🔒 password required
Public scorecard
Open ↗
GPT-4 (gpt-4-0314) — methodology validation
Both Privacy & Confidentiality and Bias & Fairness graded E under DecodingTrust's adversarial protocols; the other three dimensions not yet assessed. Methodology-validation example — real benchmark measurements, Eticas grading; not a delivered audit.
Eticas-AI/ai-audit-methodology · layers/layer-4-validations/decodingtrust/gpt-4-full/audit-findings
Open full leaflet ↗🔒 password required
Public scorecard
Open ↗
Claire — Wiselook AI-powered talent assessment platform (Bias & Fairness + Reliability; narrative channels interim)
Independent pro-bono audit, two dimensions assessed. Bias & Fairness — dimension grade D: no gender disparity in the scores that drive candidate filtering and no detectable lexical skew in the generated reports (both grade A), but the report narrative overwhelmingly fails to register injected demeaning content and can reproduce it as candidate strengths (stereotyping subcategory grade E, systemic within its subcategory). Reliability — dimension grade E: the same seeded-content instrument, read as a construct-validity question, shows the evaluation's registration of evidence is not governed by construct relevance — injected conduct that a competent evaluator would register goes unregistered by the report narrative in most delivered sections, is converted into presented strengths in a third of them, and moves the score only below the platform's per-candidate resolution. Interim: the narrative-channel figures are validated first-increment measurements, to be double-checked against the same captured evidence and complemented as further risks and mechanisms are assessed.
Eticas-AI/ai-audit-methodology · wiselook-audit findings (Bias & Fairness + Reliability)
Open full leaflet ↗🔒 password required
Public scorecard
Open ↗