Job Description
Brahma Consulting Group is conducting this search on behalf of our client.
\n
\n
About the role
\n
We're building a next-generation low-power AI accelerator (NPU) for physical AI, robotics, drones, and edge devices where latency, power, and memory movement are everything. You'll own the compiler and model-lowering stack from scratch: taking AI models and lowering them onto brand-new silicon as efficiently as the chip allows, working hand in hand with our silicon architects.
\n
This is the most important software hire on the team and a true 0-to-1 build. You'll architect the full stack yourself early on, then hire and lead the compiler team beneath you as we scale over the next year.
\n
\n
What you'll do
\n
- \n
- Build the full compiler stack end to end, from model ingestion (graph import, operator lowering, IR) to optimized executable output for our NPU
- Work with silicon architects to schedule memory movement, integrate quantization, and partition graphs for efficiency and lowest latency
- Own graph transformations, quantization integration, code generation, and compiler diagnostics
- Hire and lead the compiler and ML systems team, and serve as technical authority for evaluating future compiler talent
\n
\n
\n
\n
\n
\n
What we're looking for
\n
- \n
- 10+ years in industry, 5+ specifically in compiler development for ML accelerators, GPUs, or DSPs
- Hands-on across the full pipeline: graph import, operator lowering, compiler IR, graph transformations, quantization, and graph partitioning
- Real depth building compiler internals, not just using existing toolchains, ideally for custom or novel silicon
- A team lead who still wants to be hands-on, excited to architect and write code, not just manage
- Builder-level familiarity with MLIR, LLVM, TVM, XLA, IREE, or Glow
\n
\n
\n
\n
\n
\n