Search

[Remote in US] AI Kernel Engineer - RISC-V Software Stack

PublishedPublished: 6/14/2022
Technology

Job Description

Overview

\n

Mentium Technologies Inc. is seeking an Embedded Software Engineer to develop and optimize high-performance compute software for our custom RISC-V-based vision AI accelerator.

\n

You will work at the intersection of embedded systems, computer architecture, and machine learning, developing high-performance compute kernels, runtime components, libraries, and developer-facing SDK tools. A key part of the role will be efficiently mapping compute-intensive workloads such as convolution, matrix multiplication, and signal-processing operations onto a multicore RISC-V SoC.

\n

The role focuses heavily on vector/SIMD execution, memory optimization, data movement, multicore parallelism, and low-level performance optimization.

\n

Prior RISC-V experience is valuable but not required. Engineers with backgrounds in ARM NEON/SVE, x86 SIMD/AVX, DSP software, GPU kernel programming, embedded performance optimization, or other low-level parallel architectures are encouraged to apply.

\n

You will collaborate closely with RTL design, system architecture, software, and machine learning teams to turn architectural capabilities into a practical, high-performance, and extensible software platform.

\n


\n

Key Responsibilities

\n

    \n
  • Develop and optimize high-performance ML and DSP compute kernels, including operations such as convolution, matrix multiplication, activation functions, pooling, image-processing primitives, and related numerical workloads
  • \n

  • Optimize computationally intensive C/C++ code for vector/SIMD execution, multicore processing, and the SoC memory hierarchy
  • \n

  • Build reusable compute libraries, runtime components, APIs, and developer-facing components for the Mentium SDK
  • \n

  • Develop efficient data-movement, memory-management, and workload-scheduling strategies
  • \n

  • Optimize the use of caches, scratchpad memories, DMA engines, and on-chip memory resources
  • \n

  • Profile workloads and identify compute, memory-bandwidth, synchronization, and system-level performance bottlenecks
  • \n

  • Perform low-level performance analysis using profiling, benchmarking, cycle measurements, and hardware/software debugging tools
  • \n

  • Integrate optimized compute kernels and runtime components into AI model deployment and inference workflows
  • \n

  • Develop functional tests, performance benchmarks, reference examples, and SDK documentation
  • \n

  • Collaborate closely with RTL and system-architecture engineers to validate hardware features and improve end-to-end system performance
  • \n

  • Contribute to the architecture and programming model of Mentium's RISC-V accelerator software stack
  • \n

  • Evaluate and adapt relevant open-source libraries, runtimes, compiler technologies, and numerical software
  • \n

  • Help define software requirements and provide feedback that influences future hardware architecture
  • \n

\n


\n

Required Qualifications

\n

    \n
  • Bachelor's degree in Electrical Engineering, Computer Engineering, Computer Science, or a related technical field, or equivalent practical experience
  • \n

  • 3+ years of combined relevant industry, graduate research, doctoral research, or applied research experience
  • \n

  • Strong programming skills in C and/or C++
  • \n

  • Experience developing or optimizing performance-critical software
  • \n

  • Experience with at least one area of low-level performance programming, such as:
  • \n

  • SIMD or vector programming
  • \n

  • DSP programming
  • \n

  • GPU kernel programming
  • \n

  • Assembly or intrinsic-based optimization
  • \n

  • Performance-critical embedded software
  • \n

  • Numerical or high-performance computing
  • \n

  • Solid understanding of computer architecture, memory systems, and parallel processing
  • \n

  • Experience with performance profiling, benchmarking, low-level debugging, or cycle-level optimization
  • \n

  • Familiarity with computational workloads such as convolution, matrix multiplication, image processing, signal processing, or other numerical kernels
  • \n

  • Ability to reason about memory access patterns, data locality, computational efficiency, and hardware utilization
  • \n

  • Ability to read hardware specifications and work effectively with hardware and RTL engineers
  • \n

  • Proficiency with Python for testing, automation, benchmarking, tooling, or application development
  • \n

  • Experience with Git and standard collaborative software-development practices
  • \n

  • Strong written and verbal communication skills
  • \n

\n


\n

Preferred Qualifications

\n

Experience in several of the following areas is valuable, but we do not expect candidates to have experience with all of them:

\n

    \n
  • RISC-V instruction-set architecture or the RISC-V Vector Extension (RVV)
  • \n

  • ARM NEON or SVE, x86 SSE/AVX, DSP vector architectures, GPUs, or other SIMD/vector processors
  • \n

  • Vector intrinsics, assembly programming, compiler intrinsics, or low-level code optimization
  • \n

  • DSP, image-processing, numerical-computing, or machine-learning kernel development
  • \n

  • Quantized inference, fixed-point arithmetic, INT8/INT16 computation, FP16/BF16, or other reduced-precision numerical formats
  • \n

  • DMA, scratchpad memory, cache hierarchies, memory bandwidth optimization, and multicore synchronization
  • \n

  • Embedded, bare-metal, real-time, or resource-constrained software development
  • \n

  • Multicore SoCs or heterogeneous compute architectures
  • \n

  • Open-source RISC-V platforms such as PULP or similar multicore/accelerator systems
  • \n

  • Machine-learning frameworks and model formats such as PyTorch, TensorFlow, TFLite, or ONNX
  • \n

  • Compiler and deployment technologies such as LLVM, MLIR, TVM, Deeploy, or related systems
  • \n

  • SDKs, runtime libraries, numerical libraries, developer tools, or reusable software APIs
  • \n

  • Hardware-software co-design, SoC development, FPGA prototyping, architectural simulation, or custom accelerator development
  • \n

  • Open-source software or research software development
  • \n

\n


\n

Why Join Mentium?

\n

At Mentium, you will work at the intersection of custom silicon, RISC-V, high-performance embedded software, and AI.

\n

You will work directly with the engineers designing the underlying hardware and play a central role in determining how developers and machine-learning workloads interact with our accelerator.

\n

Rather than simply programming an existing processor, you will have the opportunity to influence the hardware-software boundary: identifying architectural bottlenecks, developing optimized compute kernels, evaluating new programming approaches, and providing feedback that can shape future generations of the hardware.

\n

Benefits:

\n

    \n
  • Competitive compensation packages
  • \n

  • Opportunity to work on diverse, cutting-edge AI projects across a range of industries.
  • \n

  • 401(k)
  • \n

  • Flexible PTO
  • \n

  • Full PPO medical, dental, and vision insurance coverage
  • \n

\n


Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...