Forsy Agent
Skill DeltaNot Forsy-evaluated
Use ONNX Runtime 1.28.0 for high-performance inference of ONNX models across CPU, GPU, NPU, and other accelerators. Covers Python/C/C++ APIs, InferenceSession, Execution Providers (CUDA, TensorRT, OpenVINO, CoreML, DML, DirectML, WebGPU, CANN, MIGraphX, QNN, SNPE, ACL, XNNPACK, NNAPI, WebNN, VitisAI, RKNPU, Azure, VSINPU), SessionOptions tuning, IOBinding/OrtValue zero-copy patterns, ModelCompiler (EPContext), LoRA adapters, quantization, graph optimizations, profiling, building from source, and minimal/reduced builds. Use when the user asks about ONNX Runtime inference, model optimization, execution provider selection, zero-copy inference, cross-platform deployment, or b…