Build your custom chip with advanced DPU silicon IPs
License the IPs at the heart of Kalray's DPUs. RISC-V cores, NPU arrays, networking and Ultra Ethernet engines, and a hardware acceleration engine. Each IP is an evolution of technologies proven across multiple generations of Kalray silicon, engineered for first-time-right integration into your next chip.
At a Glance
15+ years of DPU innovation, available as licensable IPs
7
Modular IPs across compute, networking, control
3 nm
Experience in advanced process nodes
100+
PhD engineers · 1,000+ combined years
3+
Successive Kalray DPU silicon generations
How the IPs fit together in a DPU SoC
Kalray's silicon IPs are designed to be flexibly combined. Pick between RISC-V clusters, NPU array, networking and Ultra Ethernet transport engines, a hardware acceleration engine, and a control subsystem to match your exact needs.
Modular DPU IPs for advanced custom chip
License individually or combine them. RISC-V compute, AI acceleration, networking, Ultra Ethernet transport, and hardware acceleration engines, control subsystem.
- Kalray RISC-V Efficient Core
- RISC-V Cluster
- NPU Array
- 1600G Networking Engine
- Ultra Ethernet Transport Engine
- HW Acceleration Engine
- RISC-V Control Subsystem
Kalray RISC-V Efficient Core (KV5)
A high-efficiency RISC-V core engineered for data-intensive acceleration. 8-issue superscalar with 512 bytes per cycle of throughput, and advanced vector and SIMD capabilities.
- 8-issue superscalar in-order core, RISC-V compliant
- 512 bytes per cycle load/store
- Up to 1.6 GHz frequency
- Fused multiply-add on FP16, FP32, FP64 complex vectors
RISC-V Cluster
A system-level IP organizing Kalray KV5 cores into scalable clusters, the structural foundation for data-plane and acceleration workloads at scale.
- Built on Kalray KV5 RISC-V cores
- Cluster-level cache and memory hierarchy
- Deterministic, low-latency intra-cluster communication
- Scales horizontally for data-plane workloads
Neural Processing Unit Array
Massively scalable AI inference acceleration combining compute, memory, and networking in a unified, tile-based architecture. Designed for inference at scale.
- 2D mesh scaling across compute, memory, networking tiles
- 13 TFLOPS (FP8) per compute tile
- INT8, FP16, BF16, FP8, MXFP6, MXFP4 support
- RISC-V Vector & Matrix Extensions (RVV, AME)
- 1.6 Tbps Ethernet per networking tile · 1× HBM per HBM tile
1600G Networking Engine
Up to 1.6 Tbps of hardware-based packet processing complemented by a cluster of 64 Kalray RISC-V cores handling advanced traffic management, efficiency and performance with the flexibility of programmability.
- 1.6 Tbps · 800 Mpps
- Hardware processing for Ethernet, VirtIO, Ultra Ethernet pipes
- Virtualization and cloud networking support
- Programmable datapath on 64 Kalray RISC-V cores
- Inline cryptography: IPSec, PSP, TSS
Ultra Ethernet Transport Engine
Hardware-accelerated transport offload for next-generation AI and HPC fabrics. Ultra Ethernet 1.0 compliant, enabling seamless integration into existing AI and HPC clusters with ultra-low latency and high throughput.
- Ultra Ethernet 1.0 compliant · UEC AI and HPC profiles
- SR-IOV support
- Hardware-based UE transport (SES, PDS, CMS sublayers)
- Up to 1.6 Tbps · 400M messages per second
- OFI API compatible · MPI, NCCL, RCCL support
Hardware Acceleration Engine
A flexible, high-performance platform for integrating and orchestrating third-party hardware accelerator IPs within a SoC. Combines advanced DMA, network-on-chip, and programmable RISC-V cores to enable lookaside offloads from CPUs, fast data movement, and efficient parallel processing, without redesigning control and data path logic for every accelerator IP you add.
- Generic engine enabling integration of any 3rd-party accelerator IP
- High bandwidth and high packet-per-second data transfer
- 4 Kalray RISC-V cores for pre/post processing
- High-end multi-threaded, multi-queued DMA
RISC-V Control Subsystem
A compact, fully customizable control processor providing a turnkey solution for SoC management. Built around a high-performance KV5 core, the subsystem delivers secured system management, peripheral control, and power and clock management.
- Based on Kalray RISC-V Core (KV5)
- Advanced Programmable Interrupt Controller
- Real-time Operating System
- JTAG debug, tracepoints, syscall support
- Tightly coupled memory, L2 cache, memory protection
- LLVM toolchain and debugger
Modular DPU IPs for advanced custom chip
License individually or combine them. RISC-V compute, AI acceleration, networking, Ultra Ethernet transport, and hardware acceleration engines, control subsystem.
Kalray RISC-V Efficient Core (KV5)
A high-efficiency RISC-V core engineered for data-intensive acceleration. 8-issue superscalar with 512 bytes per cycle of throughput, and advanced vector and SIMD capabilities.
- 8-issue superscalar in-order core, RISC-V compliant
- 512 bytes per cycle load/store
- Up to 1.6 GHz frequency
- Fused multiply-add on FP16, FP32, FP64 complex vectors
RISC-V Cluster
A system-level IP organizing Kalray KV5 cores into scalable clusters, the structural foundation for data-plane and acceleration workloads at scale.
- Built on Kalray KV5 RISC-V cores
- Cluster-level cache and memory hierarchy
- Deterministic, low-latency intra-cluster communication
- Scales horizontally for data-plane workloads
Neural Processing Unit Array
Massively scalable AI inference acceleration combining compute, memory, and networking in a unified, tile-based architecture. Designed for inference at scale.
- 2D mesh scaling across compute, memory, networking tiles
- 13 TFLOPS (FP8) per compute tile
- INT8, FP16, BF16, FP8, MXFP6, MXFP4 support
- RISC-V Vector & Matrix Extensions (RVV, AME)
- 1.6 Tbps Ethernet per networking tile · 1× HBM per HBM tile
1600G Networking Engine
Up to 1.6 Tbps of hardware-based packet processing complemented by a cluster of 64 Kalray RISC-V cores handling advanced traffic management, efficiency and performance with the flexibility of programmability.
- 1.6 Tbps · 800 Mpps
- Hardware processing for Ethernet, VirtIO, Ultra Ethernet pipes
- Virtualization and cloud networking support
- Programmable datapath on 64 Kalray RISC-V cores
- Inline cryptography: IPSec, PSP, TSS
Ultra Ethernet Transport Engine
Hardware-accelerated transport offload for next-generation AI and HPC fabrics. Ultra Ethernet 1.0 compliant, enabling seamless integration into existing AI and HPC clusters with ultra-low latency and high throughput.
- Ultra Ethernet 1.0 compliant · UEC AI and HPC profiles
- SR-IOV support
- Hardware-based UE transport (SES, PDS, CMS sublayers)
- Up to 1.6 Tbps · 400M messages per second
- OFI API compatible · MPI, NCCL, RCCL support
Hardware Acceleration Engine
A flexible, high-performance platform for integrating and orchestrating third-party hardware accelerator IPs within a SoC. Combines advanced DMA, network-on-chip, and programmable RISC-V cores to enable lookaside offloads from CPUs, fast data movement, and efficient parallel processing, without redesigning control and data path logic for every accelerator IP you add.
- Generic engine enabling integration of any 3rd-party accelerator IP
- High bandwidth and high packet-per-second data transfer
- 4 Kalray RISC-V cores for pre/post processing
- High-end multi-threaded, multi-queued DMA
RISC-V Control Subsystem
A compact, fully customizable control processor providing a turnkey solution for SoC management. Built around a high-performance KV5 core, the subsystem delivers secured system management, peripheral control, and power and clock management.
- Based on Kalray RISC-V Core (KV5)
- Advanced Programmable Interrupt Controller
- Real-time Operating System
- JTAG debug, tracepoints, syscall support
- Tightly coupled memory, L2 cache, memory protection
- LLVM toolchain and debugger
Built by the team behind three generations of Kalray DPUs
Advanced IPs rooted in silicon-proven technologies
Every IP is an evolution of architectures shipped in Kalray's own DPUs. Each block is thoroughly tested and simulated using leading-edge processes. Zero-bug, on-target performance the moment it arrives.
Modular by design
Compose a full SoC from NPU arrays, networking engines, RISC-V cores, and control subsystems. The portfolio is designed to be flexibly combined, not sold as a monolithic platform.
Co-design with the team that builds Kalray DPUs
From architectural definition through tape-out, integrate Kalray IPs with engineering support from the team that designs them. Accelerate time-to-market and reduce the risk that comes with custom silicon.
European engineering excellence
Headquartered in France. 15+ years of DPU expertise, experience in advanced process nodes down to 3nm, a team of more than 100 PhD engineers with over 1,000 years of combined experience.
Trusted by organizations building their own silicon
OpenChip selected Kalray to develop its Smart Data Management Unit
OpenChip is partnering with Kalray to develop a Smart Data Management Unit for AI factory and data center deployments, leveraging the Kalray IP portfolio under an existing licensing agreement, with Kalray's design team supporting integration through to silicon.
- Chip designers building custom DPUs and IPUs
- Cloud infrastructure providers
- Sovereign silicon programs
- High-performance computing vendors
Kalray's IPs are powered by the patented MPPA architecture
Massively Parallel Processor Array, a scalable, modular SoC architecture that integrates programmable cores, memory, and interconnects into a unified, deterministic computing fabric. The architectural foundation behind every Kalray IP.
Let's build your next chip together
Talk to our IP team about your custom DPU program. We'll start with your architecture and workload, and figure out which IPs make the most sense.