Skip to content

feat(inference): add full inference-models prediction backend - #1578

Closed
leeclemnet wants to merge 5 commits into
feat/export-inferencefrom
feat/inference-models-backend
Closed

leeclemnet wants to merge 5 commits into
feat/export-inferencefrom
feat/inference-models-backend

Conversation

@leeclemnet

@leeclemnet leeclemnet commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Exported prediction currently retains RF-DETR preprocessing and postprocessing, limiting end-to-end SDK acceleration. Add opt-in RFDETR.from_export(..., backend="inference_models") to run the complete SDK pipeline for ONNX/TensorRT detection, instance segmentation, and keypoint exports.

Changes

  • Return sv.Detections or sv.KeyPoints with float32 coordinates, int64 class IDs, and writable source images. Preserve SDK masks, keypoint confidence, and covariance.
  • Document SDK numerical and selection differences; reject incompatible export settings. Backbone and two-stage keypoint packages remain outside this single-artifact API.
  • Add the optional dependency, public API tests, and a CPU integration job.

Verification

Validated 8809377 on lee-t4-dev-cu128 (Tesla T4, Python 3.12, Torch 2.14, inference-models 0.39.0).

  • Full CPU suite: 6,642 passed, 232 skipped. Two workers, CUDA hidden, and bounded OMP/MKL threads:
    CUDA_VISIBLE_DEVICES='' OMP_NUM_THREADS=2 MKL_NUM_THREADS=2 \
    UV_PROJECT_ENVIRONMENT=.venv-tests uv run --no-sync pytest src/ tests/ scripts/ \
      -n 2 -m "not gpu and not e2e_roboflow" \
      --ignore=tests/run_smoke_all_models.py --ignore=tests/legacy/test_checkpoint_compat.py \
      --cov=rfdetr --cov-report=xml --timeout=240 --durations=50
  • Focused SDK backend tests: 52 passed, 7 fixture doctests skipped. Real ONNX graphs cover masks, keypoint class layouts, covariance, result dtypes, writable source images, empty results, and unsupported settings.
  • Isolated CPU CI environment with ONNX but no SDK or ONNX Runtime: 9 passed, 50 skipped. Artifact and device errors remain independent of optional runtimes; missing-SDK errors identify the correct dependency.
  • Pretrained SegSmall and KeypointPreview: export → load → predict passed on ONNX CPU, ONNX CUDA, and TensorRT CUDA. Direct SDK parity covers boxes, masks, keypoints, both confidence fields, visibility, covariance, and restored class IDs. At threshold 0.5, each runtime returns three segmented instances and one pose; threshold 1.0 returns empty results.
  • Pretrained detection Small: ONNX CUDA and TensorRT CUDA return four detections matching direct SDK output for NumPy, PIL, path, and real CUDA float-tensor inputs. PIL/path source images are writable, with float32 coordinates and int64 class IDs. Prior ONNX CPU validation at 7486715 also passed.
  • pre-commit run --all-files: passed, including mypy. Changed Python files pass LSP checks.

The full GPU suite and authenticated live-deploy tests were not run. GPU validation covered the pretrained ONNX/TensorRT cases above. Repository test workflows do not run while this PR targets feat/export-inference; the results above are VM runs.

Performance

RF-DETR detection Small at 7486715, 512×512, FP32, Tesla T4: 20 warmups and 100 timed whole predict() calls, including sv.Detections conversion, with source capture disabled.

Runtime Existing backend inference_models
ONNX CUDA 27.56 ms 24.99 ms
TensorRT CUDA 24.38 ms 20.42 ms

Values are medians. Concurrent CPU jobs on the shared VM make these timings indicative.

SegSmall ONNX CUDA at 16d3a1e, threshold 0.5, five warmups and 20 samples: native predict() 36.4 ms, SDK predict() with dense masks 44.9 ms. Direct SDK RLE plus Supervision conversion took 53.2 ms. The adapter therefore keeps dense masks. This measured segmentation case is slower with the SDK; no segmentation or keypoint acceleration claim is made. SDK 0.39 exposes no optimization-stage metadata for these task classes.

Related issues

Depends on #1563; stacked on feat/export-inference.

@socket-security

socket-security Bot commented Sep 30, 2026 •

Copy link
Copy Markdown

Review the following changes in direct dependencies. Learn more about Socket for GitHub.

Diff Package Supply Chain
Security
Vulnerability Quality Maintenance License
Addedpypi/​inference-models@​0.39.094100100100100

View full report

@Borda Borda added the enhancement New feature or request label Sep 30, 2026
@leeclemnet
leeclemnet marked this pull request as ready for review September 30, 2026 19:11
Comment thread pyproject.toml

[project.optional-dependencies]
inference-models = [
"inference-models>=0.39,<0.40; python_version < '3.14'",

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Shall we leave the extras managed by inference-models so for example
inference-models[trt10]

@Borda

Borda commented Sep 30, 2026

Copy link
Copy Markdown
Member

turning as draft until #1563 lands so we can merge it to develop

@Borda
Borda marked this pull request as draft September 30, 2026 19:53
@leeclemnet
leeclemnet force-pushed the feat/export-inference branch from e21a3ad to 8b62581 Compare September 30, 2026 21:32
@leeclemnet

leeclemnet commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor Author

closing as #1563 is going in a different direction and model.deploy_to_roboflow() already facilitates using inference

@leeclemnet leeclemnet closed this Sep 30, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

enhancement New feature or request has conflicts

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants