Multi-Modal Forensic Analysis Platform for Detecting AI-Generated Images
Production-grade forensic analyzer powered by frequency-domain mathematics, vision AI, and 7 novel detection algorithms. Engineered for accuracy.
OriginLayer is an AI image forensic platform that determines whether an image is real or AI-generated using a multi-layered detection pipeline. It combines watermark scanning, frequency analysis, vision-language models, and local classifiers into a single unified verdict β with explainable, coordinate-level evidence.
Unlike simple binary classifiers, OriginLayer identifies which AI platform generated the image (DALL-E, Midjourney, Stable Diffusion, Firefly, Gemini) and provides spatial heatmaps showing exactly where artifacts were detected.
|
|
|
|
The detection pipeline operates in 4 cascading layers:
Image Upload
β
βΌ
βββββββββββββββββββββββββββββββ
β Layer 1: Watermark Scan β SynthID Β· C2PA Β· SD Watermark
ββββββββββββββββ¬βββββββββββββββ
β
ββββββββββββΌβββββββββββ
βΌ βΌ βΌ
ββββββββββ ββββββββββ ββββββββββ
β Gemini β β DCT β β Local β
β Vision β β Freq β βClassif.β
β LLM β βAnalysisβ β(SDXL) β
ββββββ¬ββββ ββββββ¬ββββ ββββββ¬ββββ
ββββββββββββΌβββββββββββ
βΌ
βββββββββββββββββββββββββββββββ
β Multi-Modal Fusion (DRWF) β Dynamic Reliability Weighting
ββββββββββββββββ¬βββββββββββββββ
βΌ
Forensic Report
(Verdict Β· Heatmap Β· PDF)
| Layer | Technology |
|---|---|
| Frontend | Next.js 14 (App Router), React 18, Framer Motion |
| Styling | CSS Variables, Styled JSX, Glassmorphism Design System |
| ML Backend | Python FastAPI, PyTorch, HuggingFace Transformers |
| Detection APIs | Google Gemini Vision, HuggingFace Inference API |
| Local Models | Organika/sdxl-detector (ViT classifier, 331MB) |
| Forensics | DCT/FFT Analysis, GAN Fingerprinting, Watermark Detection |
| Export | html2canvas PDF generation, localStorage history |
- Node.js v20+ β
brew install node(macOS) or nodejs.org - Python 3.9+ β
brew install python@3.11(macOS)
git clone https://github.com/aashish254/ai-image-detection.git
cd ai-image-detection
npm installcp .env.example .env.localEdit .env.local with your API keys:
| Variable | Required | Get it from |
|---|---|---|
HUGGINGFACE_API_TOKEN |
β | huggingface.co/settings/tokens |
GOOGLE_GEMINI_API_KEY |
β | makersuite.google.com |
VPT_SERVER_URL |
Auto | Default: http://localhost:8100 |
Note: Model weights (~540MB total) are hosted on HuggingFace Hub β the industry standard for ML model distribution. They are not included in the GitHub repo due to size limits.
| Model | HuggingFace Repo | Size | Purpose |
|---|---|---|---|
| SDXL Detector | Organika/sdxl-detector |
331MB | Image classifier (ViT) |
| VPT Artifact Sentinel | aashish254/vpt-artifact-sentinel-v1 |
208MB | Vision-language forensic model |
python3 -m venv .venv
source .venv/bin/activate
pip install huggingface_hub
python scripts/download_model.pypip install torch torchvision torchaudio transformers peft accelerate \
pillow fastapi "uvicorn[standard]" python-multipart safetensors \
sentencepiece huggingface_hub tokenizers psutilTerminal 1 β ML Backend:
source .venv/bin/activate && cd vpt_server && python3 server.pyTerminal 2 β Frontend:
npm run devOpen http://localhost:3000 π
| Platform | Local Classifier | VPT Model | Notes |
|---|---|---|---|
| macOS (Apple Silicon) | β CPU | VPT requires NVIDIA CUDA | |
| Windows + NVIDIA GPU | β CUDA | β Full | Best experience |
| Linux + NVIDIA GPU | β CUDA | β Full | Recommended for production |
| Any (CPU only) | β CPU | Classifier works, VPT in demo mode |
originlayer/
βββ src/
β βββ app/ # Next.js App Router
β β βββ analyze/page.tsx # Analysis page
β β βββ api/analyze/route.ts # Detection API endpoint
β β βββ features/page.tsx # Features showcase
β β βββ working-model/page.tsx # How it works
β β βββ globals.css # Theme system & variables
β β βββ page.tsx # Landing page
β β
β βββ components/ # 20+ React components
β β βββ AnalysisResults.tsx # Results dashboard
β β βββ ConfidenceGauge.tsx # SVG confidence arc
β β βββ ExplanationPanel.tsx # Forensic explanation
β β βββ SpatialHeatmap.tsx # Interactive heatmap
β β βββ GANIdentificationPanel.tsx# GAN fingerprint display
β β βββ Navbar.tsx # Theme toggle & nav
β β βββ ...
β β
β βββ lib/ # Core detection logic
β βββ detectors/
β β βββ watermark.ts # SynthID/C2PA/SD detection
β β βββ vision-llm.ts # Gemini Vision integration
β β βββ huggingface.ts # HF API classifier
β β βββ dct-analysis.ts # Frequency domain analysis
β βββ fusion.ts # Multi-detector fusion (DRWF)
β βββ gradcam-analysis.ts # Forensic report generator
β βββ gan-fingerprint.ts # GAN identification
β βββ frequency-spatial-fusion.ts # FSFN implementation
β βββ uncertainty.ts # EUQ implementation
β βββ explainable-ai.ts # XAI implementation
β
βββ vpt_server/ # Python ML backend
β βββ server.py # FastAPI server
β βββ requirements.txt
β
βββ local_models/ # ML model weights (gitignored)
βββ scripts/
β βββ download_model.py # Model weight downloader
βββ docs/ # Screenshots & diagrams
βββ public/ # Static assets
-
Watermark Cascade β Scans for embedded digital watermarks (SynthID from Google, C2PA credentials, Stable Diffusion DWT-DCT marks). If found, this provides near-definitive evidence.
-
Vision-Language Analysis β Gemini Vision LLM examines the image for semantic artifacts: unnatural lighting, impossible anatomy, texture inconsistencies, and AI-typical patterns.
-
Frequency Domain Analysis β DCT spectral analysis detects statistical anomalies in frequency coefficients that are invisible to the human eye but characteristic of generative models.
-
Local Classification β The Organika/sdxl-detector ViT model provides a fast binary classification with confidence scoring.
-
Fusion & Calibration β All detector outputs are fused using Dynamic Reliability Weighted Fusion (DRWF), with disagreement-aware confidence calibration producing the final verdict.
MIT License β see LICENSE for details.
Built with β€οΈ using Next.js, Python, and advanced ML research










