ABOUT KALRAY
Kalray is a European leader in hardware acceleration, with full-stack acceleration expertise: from silicon to complete system.
Our MPPA® (Massively Parallel Processor Array) architecture is the foundation of Kalray’s processor (30+ patent families, 15+ years of development) and acceleration cards that combine processing power, flexibility, and energy efficiency.
Our mission is to deliver open data-efficient hardware accelerators to power next generation of data-intensive, AI-driven systems and infrastructures. We offer off-the-shelf processors, acceleration cards, and specialized processor development.
With over 130 employees and presence in France and Romania, Kalray is backed by top-tier investors and publicly listed on Euronext Growth. You’ll be part of a pioneering team that is shaping the future of computing with cutting-edge processor architecture, software-defined solutions, and next-generation acceleration platforms.
We offer a fast-paced, inclusive, and collaborative environment where ambitious experienced professionals and young talents can thrive—while enjoying the positive and inspiring lanscape of the Alps or the Côte d’Azur. You can learn more about us on our website, follow us on LinkedIn.
You can learn more about us on our website, follow us on LinkedIn.
WHAT YOU ‘LL BE DOING:
As an AI Kernel Optimization Engineer, you will play a key role in optimizing high-performance compute kernels to push the limits of AI Inference on RISC-V platforms.
You will design, implement, and optimize AI compute kernels (Gen AI Large Language Model, AI Vision, CNNs, etc) and runtime components to fully exploit the underlying hardware architecture – from vector/matrix units and memory hierarchies down to the assembly level.
Your work will directly influence how efficiently AI models run on SoCs, shaping the performance of next-generation inference accelerators. You will collaborate closely with hardware architects, compiler engineers, and AI framework developers to achieve optimal hardware–software co-design.
Your main responsibilities will include:
Leading and contributing to:
- Develop, optimize, profile, and debug AI compute-intensive kernels (e.g., GEMM, attention, activations) targeting RISC-V architectures.
- Identify and resolve performance bottlenecks at the ISA, compiler, and runtime levels.
- Collaborate with hardware and architecture teams to influence design decisions and improve real-world AI performance.
- Contribute to the development and optimization of AI runtime and graph execution engines.
- Evaluate, benchmark, and optimize AI inference workloads on platforms.
- Develop performance analysis tools and automation scripts for profiling, validation, and performance optimization.
- Work with AI frameworks (e.g., vLLM, SGLang, PyTorch, TensorRT-LLM) to ensure efficient mapping to targets.
- Stay up to date with AI kernel optimization trends, emerging hardware acceleration techniques, and open-source developments.
- Share technical expertise and contribute to knowledge sharing and continuous improvement within the team.
WHAT WE ARE LOOKING FOR:
Technical skills:
- Strong background in low-level performance optimization (vectorization, memory access optimization, loop unrolling, instruction scheduling, data-tiling, etc.).
- Proficiency in C/C++ and good understanding of assembly-level optimizations (SIMD, intrinsics, compiler flags).
- Solid understanding of CPU/GPU/AI accelerator architecture (pipelines, caches, memory hierarchies, compute units).
- Experience with profiling and performance analysis tools (perf, VTune, nvprof, etc.).
- Strong knowledge of parallel programming (SIMD, multithreading, OpenMP, CUDA, or similar).
- Solid software engineering skills (version control, CI/CD, testing).
- Experience with RISC-V architectures or other custom ISAs is a plus.
Nice to have:
- Experience with AI inference workloads or libraries (e.g., BLAS, cuDNN, oneDNN, TVM, or similar).
- Familiarity with MLIR/LLVM or other compiler infrastructures.
- Contributions to open-source AI inference engines or kernel libraries.
- Understanding of NUMA architectures or heterogeneous computing.
- Experience with quantization and mixed precision inference.
Profile:
- MSc or PhD in Computer Engineering or Computer Science, or equivalent practical experience.
- 3+ years of experience in performance optimization for AI Inference (or HPC use cases).
- Passion for performance and detail-oriented mindset.
- Strong analytical and problem-solving abilities.
- Comfortable working in an international, fast-evolving startup environment.
- Leadership, collaborative, and open to cross-disciplinary work.
Of course, you might not have all of those required skills! But feel free to apply anyway and explain to us why you believe you are the right person for the job.
CONTRACT INFORMATION:
- Type of contract: Permanent contract
- Starting date: As soon as possible
- Location: Montbonnot(38)
WHAT WE CAN OFFER YOU:
- Competitive salary & performance-based RSU (free shares)
- Hybrid work model
- Additional paid leave (RTT)
- Meal vouchers (Edenred)
- Premium health coverage (Malakoff Humanis)
- Sustainable mobility incentives
- Generous paternity leave
- Monthly team activities (laser game, hiking, sailing, karaoke …) and large-scale company events
RECRUITMENT PROCESS:
- First interview with the line manager
- Second interview with the team
- Final interview with HRD or CEO
Equal Opportunity Statement
KALRAY is committed to creating a diverse and inclusive environment, and we welcome applications from individuals of all backgrounds, identities, and experiences. We do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, skin color, national origin, gender, sexual orientation, age, marital status, disability status, or any other characteristic protected by law. Should you require any accommodations or adjustments throughout the interview process and beyond, please do not hesitate to let us know. We are committed to ensuring that all candidates have an equal opportunity to showcase their abilities and succeed in our organization.