Machine-learning systems now make high-stakes decisions about people, and regulation such as the EU Artificial Intelligence Act turns their fairness from an aspiration into an evidentiary obligation. Yet the evidence the field routinely produces—parity statistics computed at the output—cannot say where bias enters a system, cannot distinguish genuine debiasing from masking, and cannot certify that bias was removed rather than displaced, because the true bias of real data is unknown. This thesis builds a fairness audit stack that addresses these gaps. It first organizes tabular algorithmic fairness around the EU AI Act’s high-risk areas, unifying detection and mitigation through a shared taxonomy of intervention loci and identifying the field’s validation gap. It then contributes, layer by layer: a zero-overhead, post-training intervention that folds closed-form concept erasure exactly into trained weights, with honestly stated limits; gauge-invariant white-box instruments that provably bound the statistical-parity gap and expose masking and the instability of bias localization; a generator of domain-realistic tabular benchmarks whose bias is known by construction at controllable type, level, and severity, whose keystone experiment validates the audit hierarchy causally; and a certification protocol that turns the stack on the auditors themselves, yielding a capability-by-specification trust map for LLM-agentic bias auditors. The stack then transfers beyond tabular data: to selection bias in the multiple-choice evaluation of video language models, reduced at negligible cost under the stated protocol; to a deployed video-language pipeline, where a permutation-calibrated audit traces grounding-stage disparities through mechanism, mitigation, and propagation; and to generative face “enhancement” in forensic use, where reference-free reliability instruments ground a policy of abstention and disclosure. Together, these results move fairness auditing from asserted compliance toward verifiable, regulation-ready evidence.
I sistemi di apprendimento automatico prendono ormai decisioni ad alto rischio sulle persone, e normative come il Regolamento europeo sull’intelligenza artificiale (AI Act) trasformano la loro equità da aspirazione a obbligo probatorio. Tuttavia, l’evidenza che il settore produce abitualmente—statistiche di parità calcolate sull’output—non può dire dove il bias entri nel sistema, non può distinguere una correzione autentica dal masking, e non può certificare che il bias sia stato rimosso anziché spostato, perché il bias reale dei dati è ignoto. Questa tesi costruisce un fairness audit stack che colma tali lacune. Dapprima organizza l’equità algoritmica su dati tabellari attorno alle aree ad alto rischio dell’AI Act, unificando rilevazione e mitigazione tramite una tassonomia condivisa dei loci di intervento e identificando il divario di validazione del settore. Contribuisce quindi, livello per livello: un intervento post-addestramento a costo di inferenza nullo, che incorpora esattamente la cancellazione di concetti in forma chiusa nei pesi della rete, con limiti dichiarati; degli strumenti white-box gauge-invarianti che delimitano in modo dimostrabile il divario di parità statistica ed espongono il masking e l’instabilità della localizzazione del bias; un generatore di benchmark tabellari realistici il cui bias è noto per costruzione, a tipo, livello e severità controllabili, il cui esperimento chiave valida causalmente la gerarchia di audit; e un protocollo di certificazione che rivolge lo stack verso gli auditor stessi, producendo una mappa di fiducia capacità-per-specifica per gli auditor agentici basati su LLM. Lo stack si trasferisce poi oltre i dati tabellari: al bias di selezione nella valutazione a scelta multipla dei modelli video-linguistici, eliminato tramite calibrazione a costo trascurabile; a una pipeline video-linguistica dispiegata, dove un audit calibrato per permutazione traccia le disparità della fase di grounding attraverso meccanismo, mitigazione e propagazione; e al “miglioramento” generativo dei volti in ambito forense, dove strumenti di affidabilità privi di riferimento fondano una politica di astensione e divulgazione. Nel complesso, questi risultati spostano l’audit di equità dalla conformità asserita verso un’evidenza verificabile e pronta per la regolamentazione.
Beyond Output Parity: The Fairness Audit Stack for High-Risk AI / Bezrukov, O.. - (2026 Sep 23).
Beyond Output Parity: The Fairness Audit Stack for High-Risk AI
BEZRUKOV, OLEKSANDR
2026-09-23
Abstract
Machine-learning systems now make high-stakes decisions about people, and regulation such as the EU Artificial Intelligence Act turns their fairness from an aspiration into an evidentiary obligation. Yet the evidence the field routinely produces—parity statistics computed at the output—cannot say where bias enters a system, cannot distinguish genuine debiasing from masking, and cannot certify that bias was removed rather than displaced, because the true bias of real data is unknown. This thesis builds a fairness audit stack that addresses these gaps. It first organizes tabular algorithmic fairness around the EU AI Act’s high-risk areas, unifying detection and mitigation through a shared taxonomy of intervention loci and identifying the field’s validation gap. It then contributes, layer by layer: a zero-overhead, post-training intervention that folds closed-form concept erasure exactly into trained weights, with honestly stated limits; gauge-invariant white-box instruments that provably bound the statistical-parity gap and expose masking and the instability of bias localization; a generator of domain-realistic tabular benchmarks whose bias is known by construction at controllable type, level, and severity, whose keystone experiment validates the audit hierarchy causally; and a certification protocol that turns the stack on the auditors themselves, yielding a capability-by-specification trust map for LLM-agentic bias auditors. The stack then transfers beyond tabular data: to selection bias in the multiple-choice evaluation of video language models, reduced at negligible cost under the stated protocol; to a deployed video-language pipeline, where a permutation-calibrated audit traces grounding-stage disparities through mechanism, mitigation, and propagation; and to generative face “enhancement” in forensic use, where reference-free reliability instruments ground a policy of abstention and disclosure. Together, these results move fairness auditing from asserted compliance toward verifiable, regulation-ready evidence.I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.


