AI Research Engineer, Inference

AI Research Engineer, Inference

Job Type:

Direct-Hire

Location

Chicago

Industry:

Trading Firm

Category:

AI Engineer

Compensation Range:

$250,000 - $300,000 Per Year

Job id:

26518

Additional Compensation Info:

Base salary plus discretionary bonus

Rich Text Widget

Job Title: AI Research Engineer – Deep Learning Inference

Location: New York, NY | London, UK

About the Opportunity

Join an industry-leading global quantitative technology organization at the forefront of machine learning innovation. We are seeking a High-Performance AI Research Engineer to drive speed and scalability across our real-time predictive modeling infrastructure. In this role, you will bridge the gap between machine learning research and hardware acceleration, building low-latency inference systems that process continuous global data streams to drive core business decisions.

Responsibilities

  • Drive performance optimizations across all aspects of large-scale model execution, including custom kernel creation, data streaming pipelines, and novel hardware integration.

  • Partner directly with machine learning researchers to co-design neural architectures optimized for extreme real-time execution.

  • Author low-level primitives and custom operations to extract maximum compute throughput from modern hardware architectures.

  • Evaluate, benchmark, and deploy novel hardware acceleration tech, ranging from off-the-shelf accelerators to custom specialized silicon.

  • Formulate and execute engineering initiatives that address complex, non-obvious bottlenecks in ultra-low-latency deep learning inference.

Requirements (Must-Have)

  • At least two years of hands-on experience engineering production-grade deep learning systems within any complex domain (such as robotics, computer vision, audio, NLP, physics, or recommender platforms).

  • Strong lower-level engineering foundation, including experience writing custom compute kernels (e.g., CUDA, Triton, Pallas, or CuTe DSLs).

  • Proficiency with framework compilation internals and low-level runtime environments (PyTorch, JAX, XLA, or CUDA Graphs).

  • Practical exposure to hardware acceleration tech, such as FPGAs, ASICs, or specialized AI processors.

  • Proven ability to adapt algorithms and technical concepts across different domain applications.

Preferred Qualifications

  • Experience optimizing or serving Large Language Models (LLMs) and foundation architectures.

  • Note: Prior background in quantitative finance or trading is explicitly NOT required.

Compensation & Benefits

  • Highly competitive base salary, performance-based bonus incentive, and premium health/wellness benefits package. Equal Opportunity Employer.

Apply Now
Apply Now
Share this Job
SCHEMA MARKUP ( This text will only show on the editor. )
Back to Job Search Back to Job Search