Voice Cloning and Financial Fraud: Examining the Evidentiary Challenges under Bharatiya Sakshya Adhiniyam, 2024: AUTHOR: NIKITA Patidar

.The rapid proliferation of Generative Artificial Intelligence (GenAI) and deep-learning neural networks has fundamentally transformed audio synthesis, enabling hyper-realistic AI voice cloning. While this technological leap unlocks unprecedented capabilities across dynamic commercial applications, it simultaneously equips cybercriminals with a formidable mechanism for executing financial fraud, corporate impersonation and social engineering attacks. India’s statutory transition from the Indian Evidence Act 1872 to the Bharatiya Sakshya Adhiniyam 2023 (BSA) aimed to overhaul the evidentiary landscape to meet the demands of the digital era

ARTICLE

NIKITA Patidar

9/30/20267 min read

Abstract

The rapid proliferation of Generative Artificial Intelligence (GenAI) and deep-learning neural networks has fundamentally transformed audio synthesis, enabling hyper-realistic AI voice cloning. While this technological leap unlocks unprecedented capabilities across dynamic commercial applications, it simultaneously equips cybercriminals with a formidable mechanism for executing financial fraud, corporate impersonation and social engineering attacks. India’s statutory transition from the Indian Evidence Act 1872 to the Bharatiya Sakshya Adhiniyam 2023 (BSA) aimed to overhaul the evidentiary landscape to meet the demands of the digital era. This research paper critically evaluates the evidentiary challenges associated with proving AI voice cloning in financial fraud cases under the newly established framework of the BSA. It examines the operational mechanics of Section 63 (admissibility of electronic records), Section 57 (primary evidence) and Section 39 (expert opinion) of the BSA, contrasting them against the technical realities of generative audio. By analyzing landmark historical precedents alongside recent judicial rulings, including Pune Bar Association v Union of India, this study identifies structural lacunae in forensic voice spectrography, chain-of-custody protocols and statutory certification formats (Part A and Part B of the BSA Schedule). Ultimately, the paper proposes a dual legal and technological framework to enhance forensic capability and align judicial standards with AI-driven cybercrime vectors.

Keywords

· Voice Cloning & Generative AI

· Financial Cybercrime & Vishing

· Bharatiya Sakshya Adhiniyam 2023

· Section 63 Admissibility of Electronic Records

· Audio Spectrography & Forensic Evidence

1. Introduction

The landscape of financial crime in India is undergoing a paradigm shift. Traditional cyber enabled financial fraud, historically characterized by phishing emails, SMS spam and One-Time Password (OTP) harvesting, has rapidly evolved into AI-assisted, highly targeted attacks. Deepfake audio technology, commonly referred to as voice cloning, allows malicious actors to synthesize human speech with remarkable fidelity using short audio samples harvested from public digital footprints, social media videos or intercepted calls. By replicating acoustic properties such as fundamental frequency (pitch), speech cadence, micro-tremors and regional accent variations, fraudsters execute sophisticated vishing schemes. These include urgent family-in-distress extortion, executive impersonation (CEO fraud) and the fraudulent authorization of high-value electronic fund transfers.

From a procedural and evidentiary perspective, the proliferation of synthesized audio strikes at the core of judicial fact-finding: the determination of authorship, authenticity and record integrity. For over a century, Indian criminal and civil courts evaluated documentary and audio evidence under the framework of the Indian Evidence Act 1872. However, the enactment of the Bharatiya Sakshya Adhiniyam 2023 (BSA), which officially replaced the legacy statute on 1 July 2024, reconstructed the legal rules governing electronic records.

While the BSA introduces modernized definitions under Section 2(1)(e) and attempts to simplify electronic evidence submission via Section 63, the statute encounters profound technical friction when applied to generative audio. Unlike conventional tape recordings or digital audio files manipulated through physical editing (splicing, cutting, or phase shifting), AI voice cloning creates pristine, continuous wave files with zero structural discontinuities. Consequently, the legal system faces a dual challenge: ensuring procedural compliance through statutory certification while simultaneously verifying the substantive authenticity of digital voice recordings in an era where seeing and hearing, is no longer believing.

2. Statutory Framework under Bharatiya Sakshya Adhiniyam 2023

The legal treatment of synthesized audio within judicial proceedings relies on a tripartite statutory foundation under the BSA, read alongside relevant provisions of the Information Technology Act 2000.

A. Classification of Electronic Records (Section 2(1)(e) & Section 57)

Under Section 2(1)(e) of the BSA, an "electronic record" is defined expansively to include data, record or data generated, image or sound stored, received or sent in an electronic form. Section 57 introduces a significant modernization by treating electronic records created or stored in proper custody as primary evidence. However, when an audio file is extracted from a mobile device, server or cloud storage onto secondary media (such as a hard drive, flash drive or CD-ROM) for presentation in court, it transitions to secondary electronic evidence, bringing Section 63 into operation.

B. The Mandatory Certification Architecture of Section 63

Section 63 of the BSA replaces Section 65B of the Indian Evidence Act 1872. It establishes a complete code governing the admissibility of secondary electronic records. To satisfy Section 63, four basic conditions regarding device functionality, regular input of data and proper system operations during the material time must be met.

Furthermore, Section 63(4) mandates that any secondary electronic record must be accompanied by a statutory certificate as set out in the Schedule to the BSA:

· Part A of the Schedule: Requires the custodian of the computer or communication device to certify details such as device particulars, hash values (SHA-256 or MD5) and the integrity of the extraction process.

· Part B of the Schedule: Requires a declaration by an expert certifying the technical operation and system integrity.

C. Judicial Expert Opinion (Section 39 & IT Act Section 79A)

Section 39 of the BSA regulates the relevancy of expert opinions regarding digital signatures, electronic signatures and electronic records. In complex cybercrime prosecutions, courts rely on notified Examiners of Electronic Evidence under Section 79A of the Information Technology Act 2000. However, the current notification mechanism faces practical constraints, as standard state Forensic Science Laboratories (FSLs) lack deep-learning detection capabilities specifically trained on AI voice models.

3. Historical Precedents and Contemporary Judicial Evolution

To assess the efficacy of the BSA in handling AI voice cloning, earlier judicial precedents governing tape recordings and electronic evidence must be evaluated to determine their ongoing relevance.

A. S Partap Singh v State of Punjab

· Principle Laid Down: In S Partap Singh v State of Punjab, the Supreme Court held that tape-recorded conversations constitute primary electronic/documentary evidence, provided that the context of the conversation is relevant and the identity of the speaker is clearly established.

· Current Relevance: Partially Relevant. While the foundational requirement of logical relevance remains intact under the BSA, the baseline assumption that a recorded voice undeniably links to the physical individual is invalidated by AI voice synthesis engines capable of perfect vocal emulation.

B. RM Malkani v State of Maharashtra

· Principle Laid Down: In RM Malkani v State of Maharashtra, the Supreme Court established the classic tri-fold test for the admissibility of audio tape recordings:

1. The voice of the speaker must be identified beyond reasonable doubt.

2. The accuracy and completeness of the tape-recorded statement must be proved.

3. The possibility of tampering, erasure, or interpolation must be excluded entirely.

· Current Relevance: Highly Relevant but Harder to Satisfy. RM Malkani v State of Maharashtra remains the standard test for evaluating audio integrity in Indian jurisprudence. However, satisfying the third criterion, excluding tampering, is almost impossible using traditional forensic methods when dealing with deepfake audio. Generative Neural Networks (GNNs) construct audio frame-by-frame without creating physical cuts, background noise shifts or spectral gaps.

C. Anvar PV v PK Basheer & Arjun Panditrao Khotkar v Kailash Kushanrao Gorantyal

· Principle Laid Down: In Anvar PV v PK Basheer, as subsequently reaffirmed by a Three-Judge Bench in Arjun Panditrao Khotkar v Kailash Kushanrao Gorantyal, the Supreme Court settled the legal debate under legacy law, ruling that secondary electronic evidence is strictly inadmissible without the statutory certificate. Oral evidence or general proof cannot substitute for the mandatory written certificate.

· Current Relevance: Directly Applied under BSA Section 63. The principles affirmed in Arjun Panditrao Khotkar v Kailash Kushanrao Gorantyal apply directly to Section 63 of the BSA.

D. Pune Bar Association v Union of India

· Principle Laid Down: In the recent decision of Pune Bar Association v Union of India, the Supreme Court reiterated that the statutory certification procedure under Section 63 of the BSA is non-negotiable and mandatory for establishing the chain of custody and authenticity of digital records.

· Current Relevance: Directly Applicable. This precedent confirms that even in complex AI-driven financial crimes, procedural compliance via Section 63 certificates remains a mandatory threshold before courts can evaluate the substantive truth of an audio recording.

4. Evidentiary and Forensic Challenges in Voice Cloning Fraud

A. Traditional Audio Tampering vs. AI Voice Cloning

o Physical & Spectral Anomalies: Traditional audio shows clear splicing, sudden room-tone drops and pitch shifts. AI voice cloning creates seamless acoustic continuity and perfect replicates pitch, tone and formant frequencies.

o Metadata & Origin: Traditional edits leave software traces in files headers and link to physical devices. AI tools generate clean/spoofed metadata and use distributed networks to hide origin.

o Certification Standard: Traditional audio needs a basic non-tampering declaration, whereas AI audio requires advanced deep-learning analysis to detect.

B. The Forensic Verification Gap

Standard forensic methods like formant spectrography fail against modern generative AI models (e.g., WaveNet, Tacotron 2), leading to false positives. Verifying synthetic audio now requires specialized deep-learning models that analyze phase alignment and high-frequency synthetic artifacts.

C. Practical and Legal Challenges

o FSL Limitations: Forensic laboratories relying on old spectrographic tools frequently fail to identify deepfake audio.

o Metadata Spoofing: Vishing fraudsters Spoofing SIP headers and messaging metadata, disrupting the legal chain of custody.

o Certification Burden: Forcing fraud victims to obtain private expert certificates (Part B of BSA Schedule) causes procedural delays in freezing stolen assets.

o Standard of Proof: Under BNS 2023, defences arguments about open-source AI cloning creates “reasonable doubt” unless courts have mathematically verifiable synthetic.

5. Conclusion

The Bharatiya Sakshya Adhiniyam 2023 provides a modernized statutory frame for electronic records, yet the emergence of Generative AI voice cloning presents significant procedural and technical challenges. While Section 63 establishes clear rules for procedural admissibility, legal procedure alone cannot resolve the technical problem of distinguishing human speech from AI-generated audio.

To ensure that the BSA effectively addresses deepfake-enabled financial fraud, the following legal, technical and institutional reforms are recommended:

1. Notification of Advanced AI Forensic Examiners: The Central Government must issue notifications under Section 79A of the Information Technology Act 2000 to accredit specialized digital forensic units equipped with deep-learning detection tools capable of calculating synthetic audio probability scores.

2. Updating Judicial Interpretation of RM Malkani v State of Maharashtra: Higher courts should provide clear guidelines expanding the RM Malkani v State of Maharashtra test. In cases involving audio evidence, courts should require a technical verification report addressing synthetic audio markers alongside the mandatory Section 63 certificate.

3. Standardized Protocols for Financial Intermediaries: Banking institutions and telecom operators should adopt cryptographic verification protocols, such as signed audio stream logs and secure out-of-band transaction confirmations, to maintain an unbroken chain of custody for sensitive financial authorizations.

4. Streamlining the BSA Certification Process: To prevent Part B of the BSA Schedule from becoming an undue procedural barrier for individual fraud victims, courts should adopt a balanced approach: allow prima facie admissibility based on a valid Part A certificate, reserving mandatory expert examination for cases where the authenticity of the recording is specifically contested by the defense.

References

  1. S Partap Singh v State of Punjab AIR 1964 SC 72.

  2. RM Malkani v State of Maharashtra (1973) 1 SCC 471, 476.

  3. Anvar PV v PK Basheer (2014) 10 SCC 473, 480.

  4. Arjun Panditrao Khotkar v Kailash Kushanrao Gorantyal (2020) 7 SCC 1, 24.

  5. Pune Bar Association v Union of India 2024 INSC 582.

  6. Bharatiya Sakshya Adhiniyam 2023, s 2(1)(e).

  7. Information Technology Act 2000, s 79A.