ONNX Runtime

B
B tier on Deep Learning SoftwareScore 7.6 · #1 of 35
Android app
Yes
Free plan
Yes
Runs on
Android, iOS, Linux, Mac, self-hosted, Web, Windows
onnxruntime.ai
The ONNX Runtime homepage

Summary

ONNX Runtime is a free, cross-platform engine for machine-learning inference and training within existing technology stacks. It can run models from frameworks including PyTorch, TensorFlow/Keras, TFLite, and scikit-learn. Its inference tools optimize latency, throughput, memory use, and package size through graph optimization, accelerator-aware graph partitioning, and optimized computation kernels. The Execution Providers framework connects ONNX models to hardware-specific acceleration for CPUs, GPUs, FPGAs, and specialized NPUs. Listed providers include NVIDIA CUDA and TensorRT, Intel OpenVINO, DirectML, Qualcomm QNN, Android NNAPI, Apple CoreML, and WebGPU. Deployment options include cloud servers, edge and mobile devices, and browsers; ONNX Runtime Web supports browser models, while ONNX Runtime Mobile supports Android and iOS applications. Supported languages include Python, C#, C++, Java, JavaScript, and Rust. Developers can create smaller web or mobile packages by including only the operators and opsets their models require. It is available for Android, iOS, Linux, macOS, self-hosted environments, web, and Windows.

Who it is for

ONNX Runtime suits developers deploying machine-learning models across varied devices and hardware accelerators. It may also suit teams seeking local training or browser and mobile inference.

What is good

  • Runs models from several machine-learning frameworks.
  • Execution Providers connect models to diverse hardware accelerators.
  • Supports browser, mobile, edge, and cloud deployment.
  • Custom builds can include only required operators and opsets.

What to know first

  • Nightly builds have limited support and are discouraged for production.
  • DirectML is in sustained engineering; new Windows projects are advised to use WinML.
  • Models from untrusted sources may consume excessive memory or compute.

Everything Xiaomi review

ONNX Runtime: the full review

ONNX Runtime offers model execution and optimization across a broad range of deployment targets. Developers should heed the guidance on nightly builds, DirectML, and inspecting untrusted models.

Overview

ONNX Runtime is a free, open-source engine for running and optimizing machine-learning models inside existing applications and services. It suits developers who need to deploy models across different hardware and environments; it is less a model-building framework than a flexible execution layer.

Its broad framework and hardware support is the central draw, but deployment choices need care: nightly builds are not production-ready, and Microsoft recommends WinML rather than DirectML for new Windows projects.

Key features

Model compatibility and execution

ONNX Runtime runs models originating in PyTorch, TensorFlow and Keras, TFLite, scikit-learn, Hugging Face, and other frameworks. It supports ONNX and ORT formats, applying graph optimizations, optimized computation kernels, and graph partitioning to use available accelerators. In practice, that makes it useful when teams want to run models from different development ecosystems through a common runtime rather than bind an application to one training framework.

Optimization targets inference latency, throughput, memory use, and binary size. The trade-off is that a broad runtime is not automatically a small one: if prebuilt web or mobile packages are too large, developers can create custom builds containing only the operators and opsets their models require.

Hardware and deployment

The Execution Providers framework connects ONNX models with hardware-specific libraries across CPUs, GPUs, FPGAs, and specialized NPUs. Integrations include NVIDIA CUDA and TensorRT, Intel OpenVINO, AMD MIGraphX, Windows DirectML, Qualcomm QNN, Android NNAPI, Apple CoreML, and WebGPU, among others. This breadth gives teams room to target different accelerators, though they still need to choose and validate the appropriate provider for each deployment.

Deployment spans cloud servers, edge and mobile devices, and browsers. ONNX Runtime Web supports browser inference, while ONNX Runtime Mobile serves Android and iOS applications. The ecosystem also includes integrations with Azure Machine Learning, Azure Custom Vision, Azure SQL Edge, Azure Synapse Analytics, ML.NET, and NVIDIA Triton Inference Server.

Generative AI and training

The runtime can deploy text, image, and audio models, including Llama, Mistral, Phi, Stable Diffusion, and Whisper. Running inference on-device can keep that work private and reduce costs. ONNX Runtime also supports on-device training and says it can reduce costs for large-model training, making it relevant beyond serving predictions alone.

Security and project support

Models from untrusted sources can consume excessive memory or compute resources, so they should be inspected and tested safely before deployment. The project accepts non-trivial vulnerability reports through GitHub Security Advisories and coordinates fixes and disclosure. Documentation questions are directed to issue filing; users can also report bugs, suggest features, and contribute pull requests on GitHub.

Pricing

ONNX Runtime is free and open source under the MIT license. The Open source plan costs 0.00 USD per free and provides a cross-platform runtime. There are no paid tiers described; it is a strong fit for developers who want a no-cost runtime rather than a hosted service or a commercial support package.

Platforms

Supported platforms are Android, iOS, Linux, macOS, self-hosted deployments, web, and Windows. The listed languages include Python, C, C++, C#, Java, JavaScript, TypeScript, Kotlin, and Objective-C; the project also lists Rust support among other languages.

For Windows, the DirectML execution provider is in sustained engineering, and new projects are advised to use WinML. Nightly builds are available for testing but have limited support and are strongly discouraged for production workloads.

Who it's for

ONNX Runtime is best suited to developers integrating trained models into products or services that span frameworks, devices, or accelerator types. Its mobile and browser runtimes, provider integrations, and custom package option are particularly relevant when deployment constraints matter as much as model choice. Teams looking for a single framework for model development, or those that need a hosted product with a stated support plan, should consider whether a runtime is the right category of tool.

Pros and cons

  • Broad framework reach: Models from several major frameworks can run through one execution layer, easing cross-stack deployment.
  • Flexible acceleration: Providers cover a wide range of vendors and device types, including mobile and browser targets.
  • Deployment tuning: Inference optimizations address latency, throughput, memory, and binary size, and custom builds can trim web or mobile packages.
  • Production caution: Nightly builds have limited support and should not be used for production workloads.
  • Windows caveat: DirectML is in sustained engineering, so new Windows projects are pointed to WinML instead.
  • Model safety burden: Untrusted models may exhaust compute or memory, making inspection and safe testing important before use.

Alternatives

PyTorch is another free option with Android, iOS, Linux, macOS, self-hosted, and Windows support; choose it when you want a framework rather than an execution runtime.

TensorFlow is free and spans web and mobile as well as desktop and self-hosted platforms. It is a reasonable alternative when you want to work within TensorFlow's own framework ecosystem.

Apache TVM is a free, open-source option under the Apache License 2.0, with broad platform coverage including web. Consider it when its compiler-oriented approach is a better fit for your deployment work.

MATLAB Grader is free with a MATLAB license current under maintenance; LMS integration requires a qualifying academic license. It is aimed at a different need, so choose it when MATLAB-based grading is the task.

MegEngine is a free open-source framework with packages for several desktop systems and Android. It may suit teams seeking a framework rather than a model runtime.

Keras is a free alternative for Linux, macOS, and Windows. Choose it when those platforms and its framework role match your needs.

Trackio is free and runs across API, desktop, self-hosted, and web environments. Consider it for its library and Hugging Face hosting rather than model execution.

PaddlePaddle is a free option with Linux, API, self-hosted, and Windows support; it is another alternative to assess for framework needs.

Browse the wider Deep Learning Software category to compare tools by role.

Verdict

Choose ONNX Runtime when you need a free, cross-platform way to execute and optimize models from varied frameworks across CPUs, accelerators, mobile devices, and browsers. Its main advantage is deployment flexibility; its main caution is that provider and platform choices require attention, especially for Windows and production builds. Look elsewhere if you need a model-development framework or a managed service rather than a runtime.

ONNX Runtime plans and pricing

All plans
Open source Free MIT license · cross-platform runtime github.com · 1 Oct 2026

Compared on deep learning software

Free plan
Yesonnxruntime.ai
Training mode
localonnxruntime.ai
Deployment targets
multipleonnxruntime.ai
GPU acceleration
Yesonnxruntime.ai
Supported languages
Python, C, C++, C#, Java, JavaScript, TypeScript, Kotlin, Objective-Connxruntime.ai
Model formats
ONNX, ORTonnxruntime.ai

Facts

Purpose
ONNX Runtime is a production-grade engine for accelerating machine-learning training and inference in existing technology stacks.onnxruntime.ai · 1 Oct 2026
Model frameworks
Inference supports models from PyTorch, Hugging Face, and TensorFlow across different software and hardware stacks.onnxruntime.ai · 1 Oct 2026
Performance
It provides optimizations for inference latency, throughput, memory utilization, and binary size.onnxruntime.ai · 1 Oct 2026
Hardware acceleration
Its extensible Execution Providers framework lets ONNX models use hardware-specific acceleration libraries across CPUs, GPUs, FPGAs, and specialized NPUs.onnxruntime.ai · 1 Oct 2026
Provider integrations
Listed providers include NVIDIA CUDA and TensorRT, Intel OpenVINO, Windows DirectML, Qualcomm QNN, Android NNAPI, Apple CoreML, Azure, and WebGPU.onnxruntime.ai · 1 Oct 2026
Languages
The site lists support for Python, C#, C++, Java, JavaScript, and Rust, among other languages.onnxruntime.ai · 1 Oct 2026
Platforms
The site says ONNX Runtime runs on Linux, Windows, Mac, iOS, Android, and web browsers.onnxruntime.ai · 1 Oct 2026
Deployment
Inference is described for cloud servers, edge and mobile devices, and web browsers.onnxruntime.ai · 1 Oct 2026
Generative AI
The generative AI page describes deploying text, image, and audio models, including Llama, Mistral, Phi, Stable Diffusion, and Whisper.onnxruntime.ai · 1 Oct 2026
On-device privacy
The generative AI page says on-device models can run inference privately and save costs.onnxruntime.ai · 1 Oct 2026
Training
ONNX Runtime supports on-device training and says it can reduce costs for large-model training.onnxruntime.ai · 1 Oct 2026
Package sizing
If a prebuilt web or mobile package is too large, developers can make a custom build containing only the operators and opsets their models need.onnxruntime.ai · 1 Oct 2026
Nightly build support
The install page warns that nightly builds have limited support and advises against deploying them to production workloads.onnxruntime.ai · 1 Oct 2026
Windows guidance
The install page says DirectML is in sustained engineering and recommends WinML for new Windows projects.onnxruntime.ai · 1 Oct 2026
Maker
The site identifies Microsoft in its copyright notice; the pages reviewed do not state headquarters or a founding date.onnxruntime.ai · 1 Oct 2026
Framework support
It can run models from PyTorch, TensorFlow/Keras, TFLite, scikit-learn, and other frameworks.onnxruntime.ai · 1 Oct 2026
Inference optimization
ONNX Runtime applies graph optimizations, partitions graphs for available accelerators, and uses optimized computation kernels.onnxruntime.ai · 1 Oct 2026
Web and mobile
ONNX Runtime Web runs models in browsers, while ONNX Runtime Mobile supports Android and iOS applications.onnxruntime.ai · 1 Oct 2026
Execution providers
Execution providers include NVIDIA CUDA and TensorRT, DirectML, Intel OpenVINO, AMD MIGraphX, Qualcomm QNN, CoreML, NNAPI, and others.onnxruntime.ai · 1 Oct 2026
Integrations
The ecosystem documentation lists integrations with Azure Machine Learning, Azure Custom Vision, Azure SQL Edge, Azure Synapse Analytics, ML.NET, and NVIDIA Triton Inference Server.onnxruntime.ai · 1 Oct 2026
Security guidance
The documentation warns that models from untrusted sources may consume excessive memory or compute resources and recommends inspection and safe testing.onnxruntime.ai · 1 Oct 2026
Security reporting
The project accepts non-trivial vulnerability reports through GitHub Security Advisories and coordinates fixes and disclosure.github.com · 1 Oct 2026
Support
Documentation questions are directed to issue filing, and the project invites users to report bugs, suggest features, and submit pull requests on GitHub.onnxruntime.ai · 1 Oct 2026
Nightly builds
Nightly builds are available for testing but have limited support and are strongly discouraged for production workloads.onnxruntime.ai · 1 Oct 2026
DirectML status
The DirectML execution provider is in sustained engineering, and new Windows projects are advised to use WinML instead.onnxruntime.ai · 1 Oct 2026

Best ONNX Runtime alternatives

See all 12

Where it ranks on Everything Xiaomi

Is ONNX Runtime yours?

Claim it for free: prove the domain, then correct facts, plans and screenshots. An editor reviews every change.

Sources