TopWeb3JobsTopWeb3Jobs
AI & ML district

Machine Learning Engineer (Model Bring-Up)

Cerebras · Bengaluru, IND (Hybrid) · Full-time
◎ AI & ML🏢 Hybrid
Salary not listed

About the role

Cerebras Systems has engineered the world's largest AI processor, delivering inference speeds over 10 times faster than GPU-based cloud services. The company partners with leading AI labs, enterprises, and startups to deploy cutting-edge models at unprecedented scale. Recently, OpenAI announced a multi-year collaboration with Cerebras to operate 750 megawatts of infrastructure, fundamentally reshaping how demanding AI workloads execute in production environments.

Cerebras is at an inflection point as a company. The engineering team is growing rapidly to support dozens of model releases annually, and the culture emphasizes technical depth, research publication, and autonomy over corporate hierarchy. You'll work on one of the fastest AI supercomputers available today.

This Machine Learning Engineer role sits at the intersection of model architecture and systems optimization. You'll take models from reference implementations and make them run efficiently on Cerebras hardware, working across the full stack from model to compiler to runtime.

What you'll do

  • Bring up new models by understanding their architectures, loading and converting weights, implementing execution paths, and validating numerical correctness against reference implementations.
  • Develop and extend MLIR dialects, graph transformations, lowering passes, and hardware-specific mappings to translate models to accelerator hardware.
  • Enable model operators including attention, matrix multiplication, normalization, and positional embeddings through compiler and kernel modifications.
  • Optimize execution by improving operator fusion, tensor layouts, tiling, memory allocation, data movement, and parallel execution strategies.
  • Tune inference performance across prefill and decode phases, including KV cache management, batching strategies, and quantization approaches to enhance latency and throughput.
  • Validate model quality by investigating numerical differences and measuring accuracy impact from precision changes and compiler optimizations.
  • Profile and diagnose bottlenecks using execution traces, hardware counters, and performance tools to identify compute, memory, and communication limitations.

What you'll bring

  • Strong C++ and Python programming skills with demonstrated ability to debug complex systems.
  • Hands-on experience bringing up and debugging ML models using PyTorch or equivalent frameworks.
  • Practical knowledge of MLIR including dialects, rewrite patterns, transformation passes, and complete lowering pipelines.
  • Solid understanding of compiler fundamentals such as intermediate representations, dataflow analysis, and code generation.
  • Deep familiarity with transformer architectures, attention mechanisms, tensor operations, and numerical precision considerations.
  • Experience profiling and optimizing workloads on GPUs or other AI accelerators.
  • Proven ability to debug correctness and performance issues across model code, compiler output, kernels, and runtime layers.

Nice to have

  • Experience optimizing LLM inference including grouped-query attention, sliding-window attention, mixture of experts, KV caching, and speculative decoding.
  • Work with low-precision formats such as FP16, BF16, FP8 and understanding their accuracy and performance tradeoffs.
  • Background developing accelerator kernels or hardware-specific compiler backends.
  • Knowledge of distributed execution, model parallelism, and accelerator memory hierarchies.
  • Open-source contributions to MLIR, LLVM, inference frameworks, or related projects.

What they offer

  • Opportunity to build breakthrough AI infrastructure beyond GPU constraints.
  • Environment that supports publishing and open-sourcing cutting-edge research.
  • Role based in Bengaluru, India on a hybrid arrangement.

Pay, location & hours

Salary not listed. Based in Bengaluru, IND (Hybrid).

About Cerebras

cerebras.ai · 2 open roles in this building · Company page → · See it on the map

Apply ↗

More roles to explore

Salary not listed
Cerebras
Apply ↗

☆ Save this job

We'll e-mail you this role so you can come back to it. No account needed.

Report this job

Reports go to the TopWeb3Jobs team. Scam reports are checked first.