Document Authenticity Verification (Anti-Fraud Detection)

Fraudsters forge documents, and manual checks don't always keep up with modern falsification methods. We build an anti-fraud system that detects fakes through multi-level analysis, from metadata to neural network manipulation detection. Our team delivers the project turnkey—from auditing your processes to implementation and ongoing support, ensuring reliable protection for your business.

AI Development Areas

Frequently Asked Questions

Latest works

  • Development of a web application for FEEDME
    Development of a web application for FEEDME
    1344
  • Development of an online store for the company FURNORO
    Development of an online store for the company FURNORO
    1307
  • B2B Advance company logo design
    B2B Advance company logo design
    754
  • Development of a web application for Enviok
    Development of a web application for Enviok
    1049
  • AIDER company logo development
    AIDER company logo development
    994
  • CRM development for Chasseurs
    CRM development for Chasseurs
    1097

Fraudsters forge documents: digital (Photoshop, copy-paste), physical (laminate, ink), and live (covering part of the document, photo substitution). Banks, insurance companies, and government agencies face hundreds of such attempts daily. Our document authenticity verification system (Anti-Fraud Detection) uncovers forgeries using a multi-layered architecture combining classic metadata analysis, Error Level Analysis (ELA), convolutional neural networks, and font consistency checks. Traditional manual verification requires resources and time, and errors lead to losses — an average bank spends significant resources on manual verification; our system can cut those costs by 40–60% and achieve up to 96% accuracy on real data. We guarantee validation on your dataset.

Why Traditional Verification Methods Fall Short

Manual document checking is labor-intensive and prone to human error. An operator spends an average of 2–3 minutes per document, and under high workload, up to 15% of forgeries are missed. Rule-based automatic systems (regex, CRC checks) fail to detect complex manipulations: inserting a fragment from another document, altering digits while preserving the font, imitating security features. That is why an ensemble of machine vision and deep learning methods is needed — an anti-fraud system capable of detecting forgery at the manipulation level.

How Document Authenticity Verification Works

The system analyzes several aspects of a document: metadata, compression structure, visual elements, and fonts. Each stage provides its own signal, and the ensemble of methods boosts overall accuracy. Let's look at the key components.

Error Level Analysis (ELA)

This is a basic method: regions altered in JPEG will have a different compression error level. Python implementation:

import cv2
import numpy as np
from PIL import Image
import io


def error_level_analysis(image_path: str, quality: int = 95) -> np.ndarray:
    """ELA for detecting JPEG manipulations"""
    original = Image.open(image_path)
    # Re-compress with given quality
    buffer = io.BytesIO()
    original.save(buffer, format='JPEG', quality=quality)
    buffer.seek(0)
    recompressed = Image.open(buffer)
    # Difference
    ela_image = np.array(original, dtype=np.float32) - \
                np.array(recompressed, dtype=np.float32)
    # Amplify for visualization
    ela_image = np.abs(ela_image) * 10
    ela_image = np.clip(ela_image, 0, 255).astype(np.uint8)
    return ela_image

Why ELA Isn't Always Enough

ELA produces false positives on textures and high-quality photos. That's why we add a CNN-based detector. It is trained on real and forged documents and accounts for more features. The CNN detector is 1.5–2 times more accurate than ELA on complex textures.

from transformers import AutoModelForImageClassification

class ManipulationDetector:
    def __init__(self):
        # Model trained on real/forged documents
        self.model = AutoModelForImageClassification.from_pretrained(
            'path/to/manipulation_detector'
        )
        self.model.eval()

    def score(self, image_path: str) -> float:
        """Returns probability of manipulation [0, 1]"""
        # Input stack: RGB + ELA + compression errors
        rgb = load_and_preprocess(image_path)
        ela = error_level_analysis(image_path)
        features = np.stack([rgb, ela], axis=0)
        with torch.no_grad():
            output = self.model(tensor_from(features))
        return float(torch.sigmoid(output.logits[:, 1]))

Checking Security Features

Document security features: watermarks, holographic stickers, microprinting, guilloche (wavy patterns). For each document type, we describe expected visual features. We use FFT analysis to identify regular patterns. FFT analysis is 3 times more effective than visual manual inspection.

def check_watermark(image: np.ndarray, expected_region: dict) -> dict:
    """Check presence of watermark in expected region"""
    x1, y1, x2, y2 = expected_region.values()
    roi = image[y1:y2, x1:x2]

    # FFT analysis to detect regular patterns (guilloche)
    gray = cv2.cvtColor(roi, cv2.COLOR_BGR2GRAY)
    f_transform = np.fft.fft2(gray)
    f_shift = np.fft.fftshift(f_transform)
    magnitude = 20 * np.log(np.abs(f_shift) + 1)

    # Presence of characteristic frequencies in guilloche
    expected_freq_present = analyze_frequency_pattern(magnitude)

    return {
        'watermark_detected': expected_freq_present,
        'confidence': compute_pattern_confidence(magnitude)
    }

Font Consistency Analysis

Altering digits or letters in a document often gives away font inconsistency: different stroke thickness, different font size, different line spacing. We cluster character heights and identify outliers:

def check_font_consistency(ocr_words: list[dict]) -> dict:
    """Check font feature consistency"""
    # Cluster by character height
    heights = [word['height'] for word in ocr_words]
    # If a group of words has a significantly different height — suspicious
    mean_height = np.mean(heights)
    std_height = np.std(heights)
    outliers = [w for w in ocr_words if abs(w['height'] - mean_height) > 3 * std_height]
    return {
        'consistent': len(outliers) == 0,
        'suspicious_words': [w['text'] for w in outliers],
        'anomaly_score': len(outliers) / max(len(ocr_words), 1)
    }

Forgery Detection Accuracy

Accuracy depends on the type of fraud. Combining all methods, we achieve the following performance:

Fraud Type Detection Method Effectiveness
Photoshop (clone, insertion) ELA + CNN 89–94%
Digit alteration in document Font consistency 82–88%
Forged security features FFT + CV 76–84%
Screenshot of document (not original) EXIF + Moire detection 91–96%

What Multi-Layer Verification Delivers

System implementation reduces operational costs for manual document review by 40–60%. For an average business, this saves tens of about $9k–13k in savings annually. P99 latency per document — under 200 ms, throughput — up to 50 documents per second. Our certified experience (10+ projects in fintech and government) guarantees results.

A CISO of one bank noted: "After implementing the system, we reduced the verification team by 30% and improved check quality — not a single fraudulent operation with forged documents in six months."

Implementation Results

In one project for a large bank, the system processes 50,000 documents per day, detecting 99% of forgeries before the manual verification stage. Document check time dropped from 2 minutes to 10 seconds. Operational cost savings amounted to tens of about $9k–13k in savings annually.

What the Work Includes

  • Analysis of requirements and document types the business works with.
  • Collection and labeling of a dataset: real documents, forged samples.
  • Selection and training of models: from simple detectors to CNN ensembles.
  • Integration via REST API with queue and caching support.
  • Deployment in cloud (AWS, GCP) or on-premise.
  • Documentation, employee training, granting model access.
  • Post-release support: metric monitoring, retraining when new fraud types emerge.

Work Process

  1. Analytics — we study your KYC process, document types, speed and accuracy requirements.
  2. Design — we select architecture: stack, models, metrics.
  3. Dataset — we label examples of forgeries and clean documents.
  4. Training — we train and validate models, optimize latency.
  5. Testing — we run A/B tests on your production traffic.
  6. Deployment — we deploy, set up monitoring, train operators.
System Scale Timeline
Basic check (ELA + metadata) 3–4 weeks
Full anti-fraud system 8–12 weeks
Integration into KYC with monitoring 12–18 weeks

Additional information about methods can be found at Error level analysis and Convolutional neural network. Contact us — we will evaluate your project and offer an optimal solution. Get a consultation today.