Document Authenticity Verification (Anti-Fraud Detection)

Fraudsters forge documents: digital (Photoshop, copy-paste), physical (laminate, ink), and live (covering part of the document, photo substitution). Banks, insurance companies, and government agencies face hundreds of such attempts daily. Our document authenticity verification system (Anti-Fraud Det

AI Development Areas

Frequently Asked Questions

Latest works

  • image_web-applications_feedme_466_0.webp
    Development of a web application for FEEDME
    1285
  • image_ecommerce_furnoro_435_0.webp
    Development of an online store for the company FURNORO
    1241
  • image_logo-advance_0.webp
    B2B Advance company logo design
    696
  • image_crm_enviok_479_0.webp
    Development of a web application for Enviok
    982
  • image_logo-aider_0.webp
    AIDER company logo development
    919
  • image_crm_chasseurs_493_0.webp
    CRM development for Chasseurs
    1033

Fraudsters forge documents: digital (Photoshop, copy-paste), physical (laminate, ink), and live (covering part of the document, photo substitution). Banks, insurance companies, and government agencies face hundreds of such attempts daily. Our document authenticity verification system (Anti-Fraud Detection) uncovers forgeries using a multi-layered architecture combining classic metadata analysis, Error Level Analysis (ELA), convolutional neural networks, and font consistency checks. Traditional manual verification requires resources and time, and errors lead to losses — an average bank spends significant resources on manual verification; our system can cut those costs by 40–60% and achieve up to 96% accuracy on real data. We guarantee validation on your dataset.

Why Traditional Verification Methods Fall Short

Manual document checking is labor-intensive and prone to human error. An operator spends an average of 2–3 minutes per document, and under high workload, up to 15% of forgeries are missed. Rule-based automatic systems (regex, CRC checks) fail to detect complex manipulations: inserting a fragment from another document, altering digits while preserving the font, imitating security features. That is why an ensemble of machine vision and deep learning methods is needed — an anti-fraud system capable of detecting forgery at the manipulation level.

How Document Authenticity Verification Works

The system analyzes several aspects of a document: metadata, compression structure, visual elements, and fonts. Each stage provides its own signal, and the ensemble of methods boosts overall accuracy. Let's look at the key components.

Error Level Analysis (ELA)

This is a basic method: regions altered in JPEG will have a different compression error level. Python implementation:

import cv2 import numpy as np from PIL import Image import io def error_level_analysis(image_path: str, quality: int = 95) -> np.ndarray: """ELA for detecting JPEG manipulations""" original = Image.open(image_path) # Re-compress with given quality buffer = io.BytesIO() original.save(buffer, format='JPEG', quality=quality) buffer.seek(0) recompressed = Image.open(buffer) # Difference ela_image = np.array(original, dtype=np.float32) - \ np.array(recompressed, dtype=np.float32) # Amplify for visualization ela_image = np.abs(ela_image) * 10 ela_image = np.clip(ela_image, 0, 255).astype(np.uint8) return ela_image 

Why ELA Isn't Always Enough

ELA produces false positives on textures and high-quality photos. That's why we add a CNN-based detector. It is trained on real and forged documents and accounts for more features. The CNN detector is 1.5–2 times more accurate than ELA on complex textures.

from transformers import AutoModelForImageClassification class ManipulationDetector: def __init__(self): # Model trained on real/forged documents self.model = AutoModelForImageClassification.from_pretrained( 'path/to/manipulation_detector' ) self.model.eval() def score(self, image_path: str) -> float: """Returns probability of manipulation [0, 1]""" # Input stack: RGB + ELA + compression errors rgb = load_and_preprocess(image_path) ela = error_level_analysis(image_path) features = np.stack([rgb, ela], axis=0) with torch.no_grad(): output = self.model(tensor_from(features)) return float(torch.sigmoid(output.logits[:, 1])) 

Checking Security Features

Document security features: watermarks, holographic stickers, microprinting, guilloche (wavy patterns). For each document type, we describe expected visual features. We use FFT analysis to identify regular patterns. FFT analysis is 3 times more effective than visual manual inspection.

def check_watermark(image: np.ndarray, expected_region: dict) -> dict: """Check presence of watermark in expected region""" x1, y1, x2, y2 = expected_region.values() roi = image[y1:y2, x1:x2] # FFT analysis to detect regular patterns (guilloche) gray = cv2.cvtColor(roi, cv2.COLOR_BGR2GRAY) f_transform = np.fft.fft2(gray) f_shift = np.fft.fftshift(f_transform) magnitude = 20 * np.log(np.abs(f_shift) + 1) # Presence of characteristic frequencies in guilloche expected_freq_present = analyze_frequency_pattern(magnitude) return { 'watermark_detected': expected_freq_present, 'confidence': compute_pattern_confidence(magnitude) } 

Font Consistency Analysis

Altering digits or letters in a document often gives away font inconsistency: different stroke thickness, different font size, different line spacing. We cluster character heights and identify outliers:

def check_font_consistency(ocr_words: list[dict]) -> dict: """Check font feature consistency""" # Cluster by character height heights = [word['height'] for word in ocr_words] # If a group of words has a significantly different height — suspicious mean_height = np.mean(heights) std_height = np.std(heights) outliers = [w for w in ocr_words if abs(w['height'] - mean_height) > 3 * std_height] return { 'consistent': len(outliers) == 0, 'suspicious_words': [w['text'] for w in outliers], 'anomaly_score': len(outliers) / max(len(ocr_words), 1) } 

Forgery Detection Accuracy

Accuracy depends on the type of fraud. Combining all methods, we achieve the following performance:

Fraud Type Detection Method Effectiveness
Photoshop (clone, insertion) ELA + CNN 89–94%
Digit alteration in document Font consistency 82–88%
Forged security features FFT + CV 76–84%
Screenshot of document (not original) EXIF + Moire detection 91–96%

What Multi-Layer Verification Delivers

System implementation reduces operational costs for manual document review by 40–60%. For an average business, this saves tens of millions of rubles annually. P99 latency per document — under 200 ms, throughput — up to 50 documents per second. Our certified experience (10+ projects in fintech and government) guarantees results.

A CISO of one bank noted: "After implementing the system, we reduced the verification team by 30% and improved check quality — not a single fraudulent operation with forged documents in six months."

Implementation Results

In one project for a large bank, the system processes 50,000 documents per day, detecting 99% of forgeries before the manual verification stage. Document check time dropped from 2 minutes to 10 seconds. Operational cost savings amounted to tens of millions of rubles annually.

What the Work Includes

  • Analysis of requirements and document types the business works with.
  • Collection and labeling of a dataset: real documents, forged samples.
  • Selection and training of models: from simple detectors to CNN ensembles.
  • Integration via REST API with queue and caching support.
  • Deployment in cloud (AWS, GCP) or on-premise.
  • Documentation, employee training, granting model access.
  • Post-release support: metric monitoring, retraining when new fraud types emerge.

Work Process

  1. Analytics — we study your KYC process, document types, speed and accuracy requirements.
  2. Design — we select architecture: stack, models, metrics.
  3. Dataset — we label examples of forgeries and clean documents.
  4. Training — we train and validate models, optimize latency.
  5. Testing — we run A/B tests on your production traffic.
  6. Deployment — we deploy, set up monitoring, train operators.
System Scale Timeline
Basic check (ELA + metadata) 3–4 weeks
Full anti-fraud system 8–12 weeks
Integration into KYC with monitoring 12–18 weeks

Additional information about methods can be found at Error level analysis and Convolutional neural network. Contact us — we will evaluate your project and offer an optimal solution. Get a consultation today.