top of page
NPU Banner F.png

Neural Processing Unit (NPU) IP

ENLIGHT

ENLIGHT Classic

4/8-bit mixed-precision NPU IP

ENLIGHT Pro

Highly scalable inference NPU IP for next-gen AI applications

ENLIGHT Pro is a high-performance, scalable NPU IP designed for edge AI applications, including automotive and cameras. It supports Transformer models and delivers 4,096 INT8 MACs/cycle, with performance scalable from 8 TOPS to hundreds of TOPS. Single-, dual-, and quad-core configurations are available, along with multiple data types and tensor shape transformation operations.
 

ENLIGHT Pro incorporates a RISC-V CPU vector extension with custom instructions and supports multiple task mappings, including multiple models, data parallelism, and tensor parallelism. The ENLIGHT SDK supports widely used network formats including ONNX (PyTorch), TFLite (TensorFlow), and CFG (Darknet), providing software tools for network conversion and NPU deployment.

NPU PRO.png

ENLIGHT Pro Hardware Key Advantages

Mixed-Precision Computation (INT8, INT16, FP16)
Achieving accuracy while preserving power, performance, and area (PPA) efficiencies

Deep Neural Network (DNN)-optimized Vector Engine
Custom instructions for Softmax and local storage access & enhanced adaptability for future DNNs

Scale-out w/ Multi-core
Greater performance by parallel processing of DNN layers

Modern DNN Algorithm Support
Transformer architecture, depth-wise convolution, feature pyramid network (FPN), etc.

ENLIGHT Pro Software Key Advantages

High-level Inter-layer Optimization
Optimized layer grouping and scheduling to minimize DRAM traffic from intermediate data

DNN-layers Parallelization
Effective multi-core utilization for elevated performance & optimized core-to-core data transfer

Aggressive Quantization
Minimization of quantization loss through mixed-precision computation

Toolkit Overview

NN Converter

Converts a network file into internal network format (.enlight)​

Supports ONNX (PyTorch), TF-Lite, and CFG (Darknet)

NN Quantizer

Generates ​quantized network: float to 4-/8-bit integer

Supports per-layer quantization of activation and per-channel quantization of weight 

NN Simulator

Evaluates full precision network and quantized network​

Estimates accuracy loss due to quantization

NN Compiler

Generates NPU handling code for target architecture and network​

ENLIGHT Pro Applications

• Object detection and tracking
• Face detection and identification
• Human pose detection, gesture recognition
• Vision-based defect inspection
• Natural language interface

 

ENLIGHT Pro Deliverables

Documentation
NPU HW integration guide
NPU SW toolkit guide
Linux device driver & API manual
Technical reference manual


NPU HW
RTL (Verilog)
Example testbench
Synthesis constraints


NPU SW
Network compiler toolkit
Linux device driver

ENLIGHT Classic

4/8-bit mixed-precision NPU IP

Features a highly optimized network model compiler that reduces DRAM traffic from intermediate activation data by grouped layer partitioning and scheduling. ENLIGHT is easy to customize to different core sizes and performance for customers' targeted market applications and achieves significant efficiencies in size, power, performance, and DRAM bandwidth, based on the industry's first adoption of 4-/8-bit mixed-quantization. 

Performs various operations of deep neural networks such as convolution, pooling, and non-linear activation functions for edge computing environments. This NPU IP far surpasses alternative solutions, delivering unparalleled compute density with energy efficiency (power, performance, and area).

그림1.png

ENLIGHT Classic Hardware Key Advantages

Mixed-Precision(4/8-bit) Computation
Higher efficiency in PPAs and DRAM bandwidth​

Deep Neural Network (DNN)-optimized Vector Engine
Better adaptation to future DNN changes

Scale-out w/ Multi-core
Even higher performance by parallel processing of DNN layers

Modern DNN Algorithm Support
Depth-wise convolution, feature pyramid network (FPN), swish/mish activation, etc.

ENLIGHT Classic Software Key Advantages

High-level Inter-layer Optimization
Grouped layer partitioning and scheduling for reducing DRAM traffic from intermediate data

DNN-layers Parallelization
Efficiently utilize multi-core resources for higher performance & optimize data movements among cores

Aggressive Quantization
Maximize use of 4-bit computation capability


 

ENLIGHT Classic Toolkit Overview

NN Converter

Converts a network file into internal network format (.enlight)​

Supports ONNX (PyTorch), TF-Lite, and CFG (Darknet)

NN Quantizer

Generates ​quantized network: float to 4-/8-bit integer

Supports per-layer quantization of activation and per-channel quantization of weight 

NN Simulator

Evaluates full precision network and quantized network​

Estimates accuracy loss due to quantization

NN Compiler

Generates NPU handling code for target architecture and network​

ENLIGHT Classic Applications

• Person, vehicle, bike, traffic sign detection

• Parking lot vehicle location detection & recognition

• License plate detection & recognition

• Detection, tracking, and action recognition for surveillance

 

ENLIGHT Classic Deliverables

RTL design for synthesis

SW toolkits and device driver

User guide

Integration guide

bottom of page