
Neural Processing Unit (NPU) IP
ENLIGHT
ENLIGHT Pro
Highly scalable inference NPU IP for next-gen AI applications
ENLIGHT Pro is a high-performance, scalable NPU IP designed for edge AI applications, including automotive and cameras. It supports Transformer models and delivers 4,096 INT8 MACs/cycle, with performance scalable from 8 TOPS to hundreds of TOPS. Single-, dual-, and quad-core configurations are available, along with multiple data types and tensor shape transformation operations.
ENLIGHT Pro incorporates a RISC-V CPU vector extension with custom instructions and supports multiple task mappings, including multiple models, data parallelism, and tensor parallelism. The ENLIGHT SDK supports widely used network formats including ONNX (PyTorch), TFLite (TensorFlow), and CFG (Darknet), providing software tools for network conversion and NPU deployment.

ENLIGHT Pro Hardware Key Advantages
Mixed-Precision Computation (INT8, INT16, FP16)
Achieving accuracy while preserving power, performance, and area (PPA) efficiencies
Deep Neural Network (DNN)-optimized Vector Engine
Custom instructions for Softmax and local storage access & enhanced adaptability for future DNNs
Scale-out w/ Multi-core
Greater performance by parallel processing of DNN layers
Modern DNN Algorithm Support
Transformer architecture, depth-wise convolution, feature pyramid network (FPN), etc.
ENLIGHT Pro Software Key Advantages
High-level Inter-layer Optimization
Optimized layer grouping and scheduling to minimize DRAM traffic from intermediate data
DNN-layers Parallelization
Effective multi-core utilization for elevated performance & optimized core-to-core data transfer
Aggressive Quantization
Minimization of quantization loss through mixed-precision computation
Toolkit Overview
NN Converter
Converts a network file into internal network format (.enlight)
Supports ONNX (PyTorch), TF-Lite, and CFG (Darknet)
NN Quantizer
Generates quantized network: float to 4-/8-bit integer
Supports per-layer quantization of activation and per-channel quantization of weight
NN Simulator
Evaluates full precision network and quantized network
Estimates accuracy loss due to quantization
NN Compiler
Generates NPU handling code for target architecture and network
ENLIGHT Pro Applications
• Object detection and tracking
• Face detection and identification
• Human pose detection, gesture recognition
• Vision-based defect inspection
• Natural language interface
ENLIGHT Pro Deliverables
Documentation
NPU HW integration guide
NPU SW toolkit guide
Linux device driver & API manual
Technical reference manual
NPU HW
RTL (Verilog)
Example testbench
Synthesis constraints
NPU SW
Network compiler toolkit
Linux device driver
ENLIGHT Classic
4/8-bit mixed-precision NPU IP
Features a highly optimized network model compiler that reduces DRAM traffic from intermediate activation data by grouped layer partitioning and scheduling. ENLIGHT is easy to customize to different core sizes and performance for customers' targeted market applications and achieves significant efficiencies in size, power, performance, and DRAM bandwidth, based on the industry's first adoption of 4-/8-bit mixed-quantization.
Performs various operations of deep neural networks such as convolution, pooling, and non-linear activation functions for edge computing environments. This NPU IP far surpasses alternative solutions, delivering unparalleled compute density with energy efficiency (power, performance, and area).

ENLIGHT Classic Hardware Key Advantages
Mixed-Precision(4/8-bit) Computation
Higher efficiency in PPAs and DRAM bandwidth
Deep Neural Network (DNN)-optimized Vector Engine
Better adaptation to future DNN changes
Scale-out w/ Multi-core
Even higher performance by parallel processing of DNN layers
Modern DNN Algorithm Support
Depth-wise convolution, feature pyramid network (FPN), swish/mish activation, etc.
ENLIGHT Classic Software Key Advantages
High-level Inter-layer Optimization
Grouped layer partitioning and scheduling for reducing DRAM traffic from intermediate data
DNN-layers Parallelization
Efficiently utilize multi-core resources for higher performance & optimize data movements among cores
Aggressive Quantization
Maximize use of 4-bit computation capability
ENLIGHT Classic Toolkit Overview
NN Converter
Converts a network file into internal network format (.enlight)
Supports ONNX (PyTorch), TF-Lite, and CFG (Darknet)
NN Quantizer
Generates quantized network: float to 4-/8-bit integer
Supports per-layer quantization of activation and per-channel quantization of weight
NN Simulator
Evaluates full precision network and quantized network
Estimates accuracy loss due to quantization
NN Compiler
Generates NPU handling code for target architecture and network
ENLIGHT Classic Applications
• Person, vehicle, bike, traffic sign detection
• Parking lot vehicle location detection & recognition
• License plate detection & recognition
• Detection, tracking, and action recognition for surveillance
ENLIGHT Classic Deliverables
• RTL design for synthesis
• SW toolkits and device driver
• User guide
• Integration guide