Knowledge Base

Common Document Fraud Techniques: How Documents Are Manipulated

A structured tour of the physical, digital, and hybrid techniques attackers use to manipulate documents, with detection signals for each. Learn how Veridexa maps every technique to explainable evidence inside a multi-source fraud detection pipeline.

By Veridexa ResearchUpdated
Table of contents
  1. Introduction
  2. A Taxonomy of Document Fraud
  3. Full Counterfeits
  4. Alterations of Genuine Documents
  5. Template and Boilerplate Fraud
  6. Photograph Substitution
  7. Digital Splicing and Copy-Move
  8. Text and Field Editing
  9. Presentation Attacks and Recapture
  10. PDF-Specific Tricks
  11. Metadata Laundering
  12. Synthetic and AI-Generated Documents
  13. Impersonation and Identity Reuse
  14. Hybrid Physical-Digital Attacks
  15. Detection Signals by Technique
  16. Best Practices
  17. Veridexa Analysis
  18. Conclusion
  19. Frequently Asked Questions

Introduction

Document fraud is not a monolithic phenomenon. It is a family of techniques, each with its own tooling, cost, target, and detection signature. Understanding the taxonomy is the first step in building — and evaluating — a serious detection pipeline.

This guide surveys the most common techniques, describes what each looks like in practice, and lists the concrete signals that reveal them. It closes with how Veridexa maps every technique to explainable evidence inside a multi-source detection pipeline.

A Taxonomy of Document Fraud

  • Full counterfeits: entirely fabricated documents intended to resemble a genuine issue.
  • Alterations: genuine documents modified in one or more fields.
  • Template fraud: documents produced from a real template with fabricated data.
  • Photograph substitution: replacing the bearer's image on a real document.
  • Digital manipulation: pixel-level edits to scans or photographs.
  • Presentation attacks: recapture of screens, printouts, or masks.
  • PDF-specific manipulations: exploiting the format's structure.
  • Metadata laundering: stripping or overwriting metadata to hide provenance.
  • Synthetic generation: documents produced by generative AI systems.
  • Impersonation: using a genuine document belonging to someone else.

Full Counterfeits

Counterfeit documents are entirely fabricated to resemble a genuine issue. Quality ranges from crude novelty items to sophisticated forgeries produced by organised groups with access to security printing techniques.

  • Fonts and spacing that diverge subtly from genuine specimens.
  • Missing or malformed security features.
  • Guilloche and background patterns rendered as solid tints.
  • MRZ present but with implausible check digits or field values.
  • Photograph edges that do not integrate with the design.

Alterations of Genuine Documents

Alterations modify one or more fields on an otherwise genuine document. Because the substrate and security features are real, alterations are harder to detect than counterfeits.

  • Erasure marks or residue around altered characters.
  • Ink or toner mismatch between original and altered regions.
  • MRZ that disagrees with visually altered fields.
  • Character spacing or baseline inconsistencies in altered regions.

Template and Boilerplate Fraud

  • Genuine templates with fabricated data — common for payslips, bank statements, invoices.
  • Templates leaked or resold in fraud marketplaces.
  • Reused boilerplate across multiple fraudulent submissions.
  • Detection via layout similarity across otherwise unrelated cases.

Photograph Substitution

  • Overlay of a new photograph on a genuine card or page.
  • Digital replacement of the photograph region in a scanned image.
  • Missing or broken photo overlay security features.
  • Ghost image disagreement with the primary photograph.
  • Photograph edges that do not follow expected security patterns.

Digital Splicing and Copy-Move

  • Splicing: pasting content from another source into a document.
  • Copy-move: duplicating content from one region of the same document into another.
  • Both produce statistical anomalies detectable by image forensics.
  • Compression, resampling, and noise-level inconsistencies are typical fingerprints.

Text and Field Editing

  • Editing numeric fields — dates, amounts, IDs — in image or PDF.
  • Editing names or reference numbers to reuse a document across cases.
  • Substituted characters often show subtly different rendering.
  • PDF text-layer edits may disagree with rendered pixels.

Presentation Attacks and Recapture

Presentation attacks submit an image of an image — a photo of a printout, a capture of a screen, a masked face over a document. They are increasingly common in remote onboarding.

  • Moire patterns from screen recapture.
  • Glare and reflection patterns characteristic of printouts.
  • Compression cycles that reveal recapture history.
  • Depth and geometry signals inconsistent with a genuine document surface.

PDF-Specific Tricks

  • Incremental updates that add or hide content after signing.
  • Invisible overlays that change rendered content without touching the text layer.
  • Text layer that disagrees with rendered pixels.
  • Layered producer strings that reveal editing history.

Metadata Laundering

  • Stripping metadata entirely to hide originating tools.
  • Overwriting metadata to impersonate a legitimate producer.
  • Print-and-rescan cycles to reset metadata.
  • Sanitisation tools with characteristic producer strings.

Synthetic and AI-Generated Documents

Generative models can now produce plausible document images at scale. The detection principles remain the same — forensic, metadata, layout, and cross-evidence — augmented with generator-specific fingerprints where they exist.

  • Layout patterns that follow prompt structure rather than genuine templates.
  • Field values that violate real-world constraints (checksums, ranges, calendars).
  • Repeated stylistic tells across otherwise unrelated documents.
  • Metadata absent or inconsistent with any genuine producer.

Impersonation and Identity Reuse

  • Genuine documents used by someone other than the bearer.
  • Detection depends on biometric matching, not document authenticity.
  • Lost and stolen document registers reveal known-compromised documents.
  • Cross-application deduplication surfaces reused documents across cases.

Hybrid Physical-Digital Attacks

  • Print-then-photograph attacks that reset digital metadata while introducing physical artefacts.
  • Physical alteration followed by digital cleanup.
  • Screen recapture of a manipulated PDF displayed on a device.
  • Each stage introduces its own signals — detection benefits from stage-aware analysis.

Detection Signals by Technique

  • Counterfeits: layout deviation, security-feature absence, MRZ malformation.
  • Alterations: forensic edit signals, ink or toner mismatch, MRZ vs VIZ disagreement.
  • Templates: layout similarity across unrelated cases, boilerplate reuse.
  • Photo substitution: edge integration failure, ghost image disagreement.
  • Splicing: compression, noise, and resampling inconsistencies.
  • Text editing: font rendering deltas, PDF text-layer vs pixel disagreement.
  • Presentation attacks: moire, glare patterns, compression stacking.
  • PDF tricks: incremental update analysis, overlay detection.
  • Metadata laundering: absent or characteristic sanitiser producer strings.
  • Synthetic: constraint violations, generator fingerprints.

Best Practices

  • Never trust visual authenticity alone.
  • Combine forensic, metadata, layout, and cross-evidence signals.
  • Maintain reference specimens for genuine documents in scope.
  • Track template reuse across cases, not just per-document.
  • Route ambiguous decisions into structured manual review, not silent acceptance.
  • Preserve every submission for forensic reprocessing as detection improves.

Veridexa Analysis

Veridexa's detection layer maps directly to this taxonomy. Every technique described above has one or more corresponding signals inside the pipeline — image forensics for splicing and editing, metadata analysis for laundering and producer anomalies, MRZ validation and layout checks for counterfeit and alteration cues, and cross-evidence reasoning for impersonation and template reuse.

Every signal appears in the final report with the region or field it was derived from, so investigators can see not just that a document was flagged but which technique the evidence points to.

Conclusion

Document fraud is a broad landscape of techniques, and no single detector addresses all of them. Reliable detection requires understanding the taxonomy, instrumenting signals for each category, and combining them with explainable reasoning so that decisions are grounded in evidence rather than heuristics.

Veridexa is built on that principle: every technique maps to concrete signals, every signal is explainable, and every decision reflects the aggregate of the evidence rather than any single check.

Frequently Asked Questions

Which document fraud technique is most common today?

In digital workflows, editing fields in PDFs and photographs — dates, amounts, names — remains the most common technique because the tools are ubiquitous and the effort is low. Full counterfeits and synthetic AI-generated documents are increasingly seen at the high end.

Why isn't a clean-looking document proof of authenticity?

Modern editing tools produce visually clean output that survives casual review. Detecting manipulation reliably requires signals invisible to human inspection — forensic pixel-level analysis, metadata anomalies, layout inconsistencies, and cross-evidence disagreement.

Are AI-generated documents fundamentally different from traditional forgeries?

They differ in scale and speed rather than in the categories of manipulation. Detection principles remain similar — forensic artefacts, metadata inconsistencies, layout anomalies, and cross-evidence mismatches — with the addition of generator-specific fingerprints where they exist.

How does Veridexa map techniques to detection?

Veridexa runs image forensics, metadata analysis, layout checks, and cross-evidence reasoning together, then reports which signals contributed to the decision. Each technique described in this article has one or more detection signals inside the Veridexa pipeline, and every signal is explainable in the final report.

Detect document manipulation with Veridexa

Upload a document and receive an explainable fraud assessment that maps every detected anomaly back to the technique it most likely reflects.