An engineering service that helps chip and edge-device teams validate processor, compiler, and model-design decisions with representative machine learning workloads.
Added Aug 10, 2026
Semiconductor and edge-device teams must determine how rapidly changing machine learning models will perform on processors that are still being designed. They need scarce cross-disciplinary expertise to prototype workloads, measure performance, power, memory, and chip-area trade-offs, and convert the findings into actionable hardware and software requirements.
Offer fixed-scope architecture benchmarking engagements built around a buyer's target models, compiler stack, and processor simulator or development hardware. The service would port and optimize representative PyTorch workloads, run controlled experiments, identify bottlenecks, and deliver recommended architecture, memory-system, compiler, and model changes. Reusable benchmark harnesses and workload suites could gradually turn the consulting work into a repeatable productized service.
Model architectures are evolving faster than processor development cycles, while on-device machine learning is forcing teams to optimize across algorithms, compilers, memory, power, and hardware simultaneously. Qualcomm and Arm are both hiring for this joint exploration capability, indicating that it is strategically important and difficult to staff.
Showing 1-17 of 17 signals
Develop and scale benchmarking and workload characterization strategies to enable fast grounding-to-silicon, root-cause performance analysis, and TPU mapping optimization. Drive full-stack hardware-software co-design to optimize current and future ML accelerator architectures for business-critical production models (e.g., LLMs and embedding models).
Committed to delivering software that customers trust and confidently build their applications around. Engineers who are excited to accelerate state-of-the-art ML workloads on parallel many-core processors and scalable multi-chip systems.
Define and own the compiler architecture and technical roadmap for MTIA, including graph compilers, code generation, and optimization strategies Solve complex compiler optimization challenges spanning operator fusion, memory planning, scheduling, and efficient mapping of ML workloads to custom accelerator hardware
Within that stack, the AI kernel and optimization software development team drives the layer where architecture meets arithmetic. Our mission is performance *and* programmability at scale: hit roofline enablements on the workloads that matter, and make kernel authoring accessible enough that the whole organization can close coverage gaps without funneling every problem through a handful of experts. We do this by shipping high-performance kernel libraries with broad PyTorch operator coverage, by building the C++ and Python kernel authoring frameworks and DSL surfaces that others build on, and by writing production kernels against new architectures long before first silicon — turning hardware proposals into measured roofline evidence while the design can still change.
- **Roofline-level kernels.** GEMM and attention variants, normalization, collectives, elementwise and reduction fusions, sparse and quantized paths — implemented against novel architectural features (matrix engines, on-chip reduction fabrics, software-managed memory hierarchies) and tuned until the remaining gap to the machine's limit is explainable in a sentence. - **Numerics under precision constraints.** Low-precision formats (FP8, MX-style block-scaled types, integer quantization) where the difference between a correct scale choice and a plausible one is several decibels of signal, and where the fix has to work on silicon that has already been taped out. - **Kernel authoring frameworks.** Templateized, composable C++ kernel SDKs in the spirit of CUTLASS, Python DSLs in the spirit of Triton and CuTe, and the compiler-facing interfaces that let automated codegen reach performance that used to require a specialist. - **Pre-silicon and bring-up.** Kernels on simulators and emulators, validating architectural features and rooflines before tapeout, then first-light bring-up on real parts. - **Software mitigations for hardware realities.** Every chip ships with something you wish were different. Finding the workaround that recovers most of the lost performance — and generalizing it so nobody rediscovers it — is core to the job. Design, implement, and optimize high-performance compute and communication kernels for MTIA accelerators, taking ownership from architectural analysis through production deployment
+14 more signals