Skip to content
betoalienPublic

About

PardoX: The Hyper-Fast Data Engine

Topics

Resources

Stars

9 stars

Watchers

1 watching

Forks

Latest commit

 

History

31 Commits

Folders and files

Repository files navigation

PardoX — Hyper-Fast Data Engine

PyPI version License: MIT Powered By Rust Version

The speed of Rust. The simplicity of Python.

PardoX is a high-performance DataFrame engine for modern ETL and analytics. A Rust core — 181 C-ABI exports across 5 crates — powers identical SDKs in Python, Node.js, and PHP, with native database I/O, an ultra-fast binary format, and out-of-core processing for datasets larger than RAM.

pip install pardox

Why PardoX exists

Single-node data workloads live in an awkward gap.

On one side, pandas is simple but collapses once a dataset approaches available RAM — and its Python-object model leaves most of the CPU idle. On the other, Spark scales infinitely but drags a cluster, a JVM, and serialization overhead into problems that only ever needed one machine.

Most analytical work sits in the middle: millions to hundreds of millions of rows, on a single beefy node, where standing up a distributed cluster is pure overhead but pandas has already run out of memory.

PardoX targets exactly that gap. A Rust core maps data directly from disk or database into memory and processes it with SIMD-vectorized, multithreaded operations — no intermediate Python/JS/PHP objects, no cluster, no JVM. You keep the ergonomics of a DataFrame API in the language you already use, and the engine does the heavy lifting in native code.

It is designed for the engineer who needs Spark-class throughput on a single node without paying Spark's operational cost.


How it works

PardoX is built on three architectural decisions, each aimed at keeping the CPU busy and the data out of the host language's heap.

Zero-copy memory model. Data flows from disk or a database socket straight into memory-mapped Rust buffers (HyperBlocks). The Python, Node, or PHP SDK holds an opaque pointer, never the raw bytes — so there is no per-row object allocation, no marshalling tax, and no garbage-collector pressure in the host runtime.

SIMD + multithreading in the core. Arithmetic and comparison operations compile down to AVX2 (x86) or NEON (ARM/Apple Silicon) vector instructions and fan out across cores. Column operations that would loop row-by-row in Python run 5×–20× faster.

A native binary format. The .prdx format is laid out for sequential, memory-mappable reads, reaching ~4.6 GB/s read throughput on repeated workloads — faster than Parquet when the same data is read many times, which is the common case in iterative ETL and feature engineering.

                ┌──────────────────────────────────────────────┐
 Python / JS /  │              Thin SDK layer                  │  Opaque pointer only
 PHP call ────▶ │        (DataFrame API, your language)         │  — no raw bytes cross
                └───────────────────────┬──────────────────────┘  the FFI boundary
                                        │ C-ABI (181 exports)
                ┌───────────────────────▼──────────────────────┐
                │                  Rust core                    │
                │  HyperBlock buffers · SIMD ops · thread pool  │
                │  tokio DB drivers · .prdx reader/writer       │
                └───────────────────────┬──────────────────────┘
                                        │ memory-mapped / async I/O
                ┌───────────────────────▼──────────────────────┐
                │   Disk (.prdx) · PostgreSQL · MySQL · Mongo   │
                │   SQL Server · S3 / GCS / Azure · Arrow       │
                └──────────────────────────────────────────────┘

Where PardoX fits

PardoX is not trying to replace your whole stack. It occupies a specific niche, and it's worth being honest about where each tool wins:

Tool Best at Trade-off vs. PardoX
pandas Small-to-medium data, richest ecosystem Dies past RAM; single-threaded; object overhead
Polars Fast single-node DataFrames Excellent overlap; PardoX adds native multi-DB I/O, .prdx format, and a multi-language (Py/JS/PHP) core
DuckDB In-process analytical SQL SQL-first; PardoX is DataFrame-first with streaming DB writes and a shared core across three languages
Spark True distributed, petabyte-scale Cluster + JVM overhead; PardoX targets single-node workloads where that overhead isn't justified

Use PardoX when you're on a single node, your data is larger than comfortable for pandas, you want native database read/write without Python drivers, or you need the same engine callable from Python, Node, and PHP.

Reach for Spark when you genuinely need horizontal scale across a cluster. PardoX is deliberately a single-node engine.


Core capabilities

  • Zero-copy architecture — Rust HyperBlock buffers, no intermediate host-language objects.
  • SIMD + multithreading — AVX2/NEON vectorized ops, 5×–20× over Python loops.
  • Native database I/O — PostgreSQL, MySQL, SQL Server, MongoDB through Rust drivers. No psycopg2, no pymysql.
  • .prdx binary format — ~4.6 GB/s read throughput; out-of-core processing for datasets larger than RAM.
  • Streaming SQL — server-side cursors and COPY FROM STDIN; 150M rows / 3.8 GB streamed to PostgreSQL at ~300k rows/s with O(block) RAM.
  • GPU sort — sort_values(gpu=True), WebGPU Bitonic sort with automatic CPU fallback.
  • ML-ready — zero-copy NumPy bridge via the __array__ protocol for direct scikit-learn interop.
  • Multi-SDK — one Rust core, identical API in Python, Node.js, and PHP.
  • 30 feature gaps implemented — GroupBy, window functions, lazy pipelines, SQL-over-DataFrames, encryption, data contracts, time travel, Arrow Flight, cloud storage, and more.

Full release history in CHANGELOG.md.


Quick start

import pardox as px

# Load a CSV with the parallel Rust parser
df = px.read_csv("sales.csv")
print(f"{df.shape[0]:,} rows × {df.shape[1]} columns")

# GroupBy — executed entirely in Rust
grouped = df.groupby("state", {"revenue": "sum", "qty": "count"})

# Stream 150M rows to PostgreSQL with O(block) RAM
rows = px.write_sql_prdx(
    "sales_150m.prdx",
    "postgresql://user:pass@localhost:5432/db",
    "sales", mode="append", conflict_cols=[], batch_rows=1_000_000
)

# Out-of-core: GroupBy on a .prdx file without loading it all into RAM
result = px.prdx_groupby("sales_150m.prdx", ["region"], {"revenue": "sum"})

Install

# Python
pip install pardox

# Node.js
npm i @pardox/pardox

# PHP
composer require betoalien/pardox-php

Benchmarks

Measured against standard Python tooling on the same hardware. Reproduction scripts live in /python files.

Operation Baseline PardoX Speedup
Read CSV (1 GB) pandas ~4.2s ~0.8s 5×
Column multiply (1M rows) pandas ~0.15s ~0.02s 7.5×
PostgreSQL write (50k rows) psycopg2 ~18s ~0.6s via COPY 30×
MySQL write (50k rows) pymysql ~22s ~3s (batch) 7×
.prdx → PostgreSQL (150M rows) — ~490s (~306k rows/s) streaming, O(block) RAM

Benchmarks reflect single-node performance. Numbers vary with hardware, schema, and workload — the reproduction scripts let you verify on your own machine.


Documentation

Full documentation, guides, and API reference at pardox.io.

Section
Getting Started Installation · Quick Start · Roadmap
User Guide I/O · Databases · Mutation & Arithmetic · Aggregations · GPU · ML Integration
API Reference Full Reference · FFI Exports (181 C-ABI functions)
SDKs Python · Node.js · PHP · Database Integration
Examples Jupyter Notebooks · 640M-row benchmark scripts

Community

  • X / Twitter: @pardox_io
  • Issues & feature requests: GitHub Issues — several shipped features (including the SQL Cursor API) started as community requests.

License

MIT — see LICENSE.

Copyright © 2026 PardoX.

About

PardoX: The Hyper-Fast Data Engine

Topics

Resources

Stars

9 stars

Watchers

1 watching

Forks

Releases

Sponsor this project

Packages

Contributors

Languages