AI in Auditing: Revolution or Risk? An Empirical Analysis Using IFRS 16 as an Example
Artificial intelligence is rapidly transforming our working world—but how does it perform in highly complex, rule-bound tasks of international accounting? A graduate of Allensbach University Konstanz investigated precisely this question in his bachelor's thesis in the Business Administration and Management degree program. His empirical analysis demonstrates in a well-founded manner where the potentials of ChatGPT and Google Gemini lie, why prompt engineering makes the decisive difference, and where humans remain indispensable.
The integration of large language models (LLMs) into everyday business operations is now unstoppable. According to recent surveys, more than 41 % of German companies are already using AI-powered solutions. Given the ongoing shortage of skilled workers in the finance and auditing sectors, the technology is seen as a major source of hope for increasing efficiency.
However, accounting standards such as IFRS 16 (Leases) do not forgive mistakes. The standard regulates the accounting of leases and requires the lessee to no longer just book simple rental expenses, but to recognize a „right-of-use asset“ on the asset side and a corresponding lease liability on the liability side. This requires both precise legal subsumption (Is there a lease at all?) and complex financial-mathematical calculations across multiple periods.
Can a general, freely accessible AI master these tasks without errors?
The study design: 720 test runs on the test bench
To provide a reliable assessment of the performance of modern language models, the student developed an extremely robust empirical research design. The models ChatGPT (GPT-4o) and Google Gemini were tested using twelve practical IFRS 16 scenarios.
To rule out statistical coincidences, each issue was queried ten times per model in completely isolated chats (zero-state design). This results in an impressive total of 720 independent test runs.
The investigation was carried out in three successively building prompt stages to measure the direct influence of the phrasing on the result:
- Level 1 (Zero-Shot): The AI processes the facts completely unprepared without additional assistance.
- Level 2 (In-Context Learning): The prompt is supplemented by the full text of IFRS 16 and an ideal provisioning example (few-shot).
- Level 3 (Expert Prompting): The AI is specifically assigned the role of an auditor and guided through a structured audit framework and step-by-step „chain-of-thought“ instructions.
In line with auditing practice, a quantitative materiality threshold of a maximum deviation of 5 % from the mathematically correct target value was defined as a critical evaluation criterion.
The core findings: Light and shadow on three levels
The evaluation of the 720 test runs provides fascinating yet sobering insights for practical application.
- The qualitative recognition accuracy (Is there a lease case?)
Surprisingly, both models demonstrated high baseline accuracy on simple standard scenarios even in their unprepared state (Level 1). Google Gemini achieved an error rate of just 2.50 % in the zero-shot approach, while ChatGPT had an error rate of 8.33 %.
The problem: As soon as an issue became more linguistically or legally demanding, the models failed collectively. A particularly striking case involved a contractually agreed exchange right for a company car. According to IFRS 16, there is no lease if the Lessor a substantial right of exchange is granted. In the test case, however, the right lay with the lessee. Both models were misled by the keyword „exchange“ and incorrectly classified the case as „No lease.“ Moreover, simply feeding them the full IFRS text in stage 2 actually worsened the error rate for this special case, as the AI was unable to correctly assign the legal roles without procedural guidance.
- The quantitative calculation accuracy
At the mathematical level, the enormous value of professional prompt engineering becomes apparent:
- The aggregate total error gradually decreased from 16.40 % (Level 1) to 12.10 % (Level 2) and then to 11.50 % (Level 3).
- Simple present value calculations and subsequent valuations (as in Scenarios 1, 3, and 4) were solved correctly and consistently in Stage 3 by both models, with a deviation of exactly 0.00 %.
- Complex tasks, such as the mathematical reconstruction of an implicit interest rate (case 8), where the models in tier 1 still produced absurd mathematical hallucinations (such as negative interest rates), completely stabilized in tier 3 thanks to the structured guidance.
The catch: Despite all the optimizations, the average maximum deviation (the worst-case scenario) was still 24.80 % even at the highest prompt level 3—and thus dramatically above the materiality threshold of 5 %. Furthermore, when dealing with high complexity at Level 3, ChatGPT tended to simply refuse to perform the mathematical calculation altogether in nearly one out of every three runs (29 %) and instead resort to purely qualitative, vague plausibility checks.
The additional experiment: The devil is in the details
The student proved that general language models must not be blindly used as a „black box“ in a final additional experiment regarding the persistent error in case study 6 (company car/exchange right). They purposefully expanded the expert template from level 3 with specific references from current specialized literature on the economic substance of exchange rights.
The result was spectacular: In a direct comparison, the number of errors dropped sharply from 13 to just one in 20 runs. The success rate jumped from 35 % to 95 %. This clearly demonstrates that AI systems benefit enormously from precise, specialized-literature-based context control.
Cost-efficiency and data protection in daily audit routines
In addition to the technical quality, the paper also examines the economic component. The use of generative AI in finance is a double-edged sword. On the one hand, the costs for data-protected enterprise solutions (such as ChatGPT Business for approx. 50 USD/month) already with a monthly time saving of just under 3 working hours based on the minimum wage. However, since the creation of high-quality prompt patterns requires deep IFRS expertise, the initial time investment is high – but it pays off due to the multiple reusability of standardized templates in everyday auditing.
This is countered by significant liability and reputational risks from undetected AI errors or hallucinations. In addition, professional and data protection laws prohibit the uncritical input of sensitive, real client data into publicly accessible systems.
Conclusion: AI as a powerful co-pilot, not an autopilot
The bachelor thesis underpins a central thesis of modern business informatics: generative AI is not to be understood in the final examination as a replacement for human expertise, but as a powerful, supporting tool. It functions excellently as a „co-pilot“ that prestructures complex issues, prepares initial mathematical calculations, and quickly makes anomalies in large amounts of text visible.
However, final responsibility, the critical assessment of incomplete or ambiguous contracts, and the ultimate audit opinion must remain strictly in the hands of the qualified auditor. „Automation bias“—blind trust in elegantly generated AI outputs that are incorrect in content—currently represents the greatest risk.
We congratulate the student on this outstanding, practice-oriented research paper, which impressively demonstrates how cutting-edge technological trends are combined with sound business expertise at Allensbach University!