Policy Semantic Audit

AI-driven expense justification research platform

E1: OCR Extraction Baseline

Establish a baseline for OCR extraction accuracy by capturing all structured information from real-world receipts.

E2: Direct Policy Reasoning

Apply organisational policy directly to structured accounting data without structured reasoning.

E3: Structured CoT

Test whether structured chain-of-thought reasoning improves policy interpretation and verdict accuracy.

E4: Context-Sensitive Reasoning

Evaluate how project phase context affects the same transaction's compliance outcome.

E5: Uncertainty Handling

Test whether the model recognises ambiguity and avoids forced confident verdicts on incomplete evidence.

E6: Adversarial Stress Test

Run the same transaction against all three policy variants simultaneously to surface policy-quality issues.

Ground Truth Annotation

Annotate each receipt with a human verdict and supporting reason. The reference dataset for evaluation.

Policy Editor

Define and version the three controlled policy variants used by E2–E6.

Research Configuration

Configure the API key, models, currency treatment and saved-result controls used across all experiments.