Skip to content

Repository files navigation

OriginLayer Banner

OriginLayer

Multi-Modal Forensic Analysis Platform for Detecting AI-Generated Images

Version Next.js Python PyTorch License

Production-grade forensic analyzer powered by frequency-domain mathematics, vision AI, and 7 novel detection algorithms. Engineered for accuracy.


πŸ“Έ Screenshots

Landing Page

Landing Page

Image Upload & Forensic Scanning

Upload Page Forensic Scanning

Analysis Results & Confidence Gauge

Analysis Results Detailed Analysis

GAN Fingerprint & Spatial Heatmap

GAN Fingerprint Identification Spatial Artifact Heatmap

Regional Distribution & Forensic Explanation

Regional AI Distribution Forensic Explanation


πŸ” What is OriginLayer?

OriginLayer is an AI image forensic platform that determines whether an image is real or AI-generated using a multi-layered detection pipeline. It combines watermark scanning, frequency analysis, vision-language models, and local classifiers into a single unified verdict β€” with explainable, coordinate-level evidence.

Unlike simple binary classifiers, OriginLayer identifies which AI platform generated the image (DALL-E, Midjourney, Stable Diffusion, Firefly, Gemini) and provides spatial heatmaps showing exactly where artifacts were detected.


✨ Key Features

πŸ›‘οΈ Detection Engine

  • Cascade Watermark Detection β€” SynthID, C2PA, SD watermarks
  • GAN Fingerprint Identification β€” Platform-level attribution
  • DCT Frequency Analysis β€” Spectral anomaly detection
  • Gemini Vision LLM β€” Semantic artifact analysis
  • Local Classifier β€” Organika/sdxl-detector on CPU

🧠 Novel Research (7 Contributions)

  • FSFN β€” Frequency-Spatial Fusion Network
  • XAI β€” Explainable AI with attention maps
  • EUQ β€” Epistemic Uncertainty Quantification
  • GFI β€” GAN Fingerprint Identification
  • DACC β€” Dynamic Accuracy Cross-Calibration
  • SAM β€” Spatial Artifact Mapping
  • DRWF β€” Dynamic Reliability Weighted Fusion

🎨 Professional UI

  • Premium dark/light mode with animated toggle
  • Glassmorphism design system
  • Framer Motion micro-animations
  • Responsive across all devices

πŸ“Š Forensic Output

  • Coordinate-level artifact detection
  • Interactive spatial heatmaps
  • Confidence gauge with calibration
  • Downloadable PDF forensic reports
  • Analysis history with local storage

πŸ—οΈ System Architecture

System Architecture

The detection pipeline operates in 4 cascading layers:

Image Upload
    β”‚
    β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚  Layer 1: Watermark Scan    β”‚  SynthID Β· C2PA Β· SD Watermark
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
               β”‚
    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
    β–Ό          β–Ό          β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β” β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β” β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ Gemini β”‚ β”‚  DCT   β”‚ β”‚ Local  β”‚
β”‚ Vision β”‚ β”‚  Freq  β”‚ β”‚Classif.β”‚
β”‚  LLM   β”‚ β”‚Analysisβ”‚ β”‚(SDXL)  β”‚
β””β”€β”€β”€β”€β”¬β”€β”€β”€β”˜ β””β”€β”€β”€β”€β”¬β”€β”€β”€β”˜ β””β”€β”€β”€β”€β”¬β”€β”€β”€β”˜
     β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚  Multi-Modal Fusion (DRWF)  β”‚  Dynamic Reliability Weighting
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
               β–Ό
       Forensic Report
  (Verdict Β· Heatmap Β· PDF)

πŸ› οΈ Tech Stack

Layer Technology
Frontend Next.js 14 (App Router), React 18, Framer Motion
Styling CSS Variables, Styled JSX, Glassmorphism Design System
ML Backend Python FastAPI, PyTorch, HuggingFace Transformers
Detection APIs Google Gemini Vision, HuggingFace Inference API
Local Models Organika/sdxl-detector (ViT classifier, 331MB)
Forensics DCT/FFT Analysis, GAN Fingerprinting, Watermark Detection
Export html2canvas PDF generation, localStorage history

πŸš€ Quick Start

Prerequisites

  • Node.js v20+ β€” brew install node (macOS) or nodejs.org
  • Python 3.9+ β€” brew install python@3.11 (macOS)

1. Clone & Install

git clone https://github.com/aashish254/ai-image-detection.git
cd ai-image-detection
npm install

2. Configure Environment

cp .env.example .env.local

Edit .env.local with your API keys:

Variable Required Get it from
HUGGINGFACE_API_TOKEN βœ… huggingface.co/settings/tokens
GOOGLE_GEMINI_API_KEY βœ… makersuite.google.com
VPT_SERVER_URL Auto Default: http://localhost:8100

3. Download ML Models

Note: Model weights (~540MB total) are hosted on HuggingFace Hub β€” the industry standard for ML model distribution. They are not included in the GitHub repo due to size limits.

Model HuggingFace Repo Size Purpose
SDXL Detector Organika/sdxl-detector 331MB Image classifier (ViT)
VPT Artifact Sentinel aashish254/vpt-artifact-sentinel-v1 208MB Vision-language forensic model
python3 -m venv .venv
source .venv/bin/activate
pip install huggingface_hub
python scripts/download_model.py

4. Install Python Dependencies

pip install torch torchvision torchaudio transformers peft accelerate \
  pillow fastapi "uvicorn[standard]" python-multipart safetensors \
  sentencepiece huggingface_hub tokenizers psutil

5. Run

Terminal 1 β€” ML Backend:

source .venv/bin/activate && cd vpt_server && python3 server.py

Terminal 2 β€” Frontend:

npm run dev

Open http://localhost:3000 πŸš€


πŸ’» Platform Compatibility

Platform Local Classifier VPT Model Notes
macOS (Apple Silicon) βœ… CPU ⚠️ Demo VPT requires NVIDIA CUDA
Windows + NVIDIA GPU βœ… CUDA βœ… Full Best experience
Linux + NVIDIA GPU βœ… CUDA βœ… Full Recommended for production
Any (CPU only) βœ… CPU ⚠️ Demo Classifier works, VPT in demo mode

πŸ“ Project Structure

originlayer/
β”œβ”€β”€ src/
β”‚   β”œβ”€β”€ app/                          # Next.js App Router
β”‚   β”‚   β”œβ”€β”€ analyze/page.tsx          # Analysis page
β”‚   β”‚   β”œβ”€β”€ api/analyze/route.ts      # Detection API endpoint
β”‚   β”‚   β”œβ”€β”€ features/page.tsx         # Features showcase
β”‚   β”‚   β”œβ”€β”€ working-model/page.tsx    # How it works
β”‚   β”‚   β”œβ”€β”€ globals.css               # Theme system & variables
β”‚   β”‚   └── page.tsx                  # Landing page
β”‚   β”‚
β”‚   β”œβ”€β”€ components/                   # 20+ React components
β”‚   β”‚   β”œβ”€β”€ AnalysisResults.tsx       # Results dashboard
β”‚   β”‚   β”œβ”€β”€ ConfidenceGauge.tsx       # SVG confidence arc
β”‚   β”‚   β”œβ”€β”€ ExplanationPanel.tsx      # Forensic explanation
β”‚   β”‚   β”œβ”€β”€ SpatialHeatmap.tsx        # Interactive heatmap
β”‚   β”‚   β”œβ”€β”€ GANIdentificationPanel.tsx# GAN fingerprint display
β”‚   β”‚   β”œβ”€β”€ Navbar.tsx                # Theme toggle & nav
β”‚   β”‚   └── ...
β”‚   β”‚
β”‚   └── lib/                          # Core detection logic
β”‚       β”œβ”€β”€ detectors/
β”‚       β”‚   β”œβ”€β”€ watermark.ts          # SynthID/C2PA/SD detection
β”‚       β”‚   β”œβ”€β”€ vision-llm.ts         # Gemini Vision integration
β”‚       β”‚   β”œβ”€β”€ huggingface.ts        # HF API classifier
β”‚       β”‚   └── dct-analysis.ts       # Frequency domain analysis
β”‚       β”œβ”€β”€ fusion.ts                 # Multi-detector fusion (DRWF)
β”‚       β”œβ”€β”€ gradcam-analysis.ts       # Forensic report generator
β”‚       β”œβ”€β”€ gan-fingerprint.ts        # GAN identification
β”‚       β”œβ”€β”€ frequency-spatial-fusion.ts # FSFN implementation
β”‚       β”œβ”€β”€ uncertainty.ts            # EUQ implementation
β”‚       └── explainable-ai.ts         # XAI implementation
β”‚
β”œβ”€β”€ vpt_server/                       # Python ML backend
β”‚   β”œβ”€β”€ server.py                     # FastAPI server
β”‚   └── requirements.txt
β”‚
β”œβ”€β”€ local_models/                     # ML model weights (gitignored)
β”œβ”€β”€ scripts/
β”‚   └── download_model.py             # Model weight downloader
β”œβ”€β”€ docs/                             # Screenshots & diagrams
└── public/                           # Static assets

πŸ”¬ How Detection Works

  1. Watermark Cascade β€” Scans for embedded digital watermarks (SynthID from Google, C2PA credentials, Stable Diffusion DWT-DCT marks). If found, this provides near-definitive evidence.

  2. Vision-Language Analysis β€” Gemini Vision LLM examines the image for semantic artifacts: unnatural lighting, impossible anatomy, texture inconsistencies, and AI-typical patterns.

  3. Frequency Domain Analysis β€” DCT spectral analysis detects statistical anomalies in frequency coefficients that are invisible to the human eye but characteristic of generative models.

  4. Local Classification β€” The Organika/sdxl-detector ViT model provides a fast binary classification with confidence scoring.

  5. Fusion & Calibration β€” All detector outputs are fused using Dynamic Reliability Weighted Fusion (DRWF), with disagreement-aware confidence calibration producing the final verdict.


πŸ“„ License

MIT License β€” see LICENSE for details.


Built with ❀️ using Next.js, Python, and advanced ML research

About

πŸ” Multi-Modal Forensic Analysis Platform β€” Detect AI-generated images using frequency-domain math, vision AI & 7 novel detection algorithms

Topics

Resources

Stars

6 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages