Jobs
Required Skills
Kubernetes
About xAI
xAI’s mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company’s mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates.
ABOUT THE ROLE:
The ML and Data Infrastructure team is responsible for building the foundational infrastructure that powers frontier AI models and truth-seeking agents—from petabyte-scale data acquisition and multimodal crawling, to web-scale search/retrieval systems, reliable high-throughput inference serving, low-level GPU/kernel optimizations, compiler/runtime innovations, and high-speed interconnect fabrics for massive clusters. In this role, you will collaborate across pre-training, multimodal, reasoning, and product teams in a fast-paced, meritocratic environment where you will tackle ambiguous, high-stakes problems with first-principles thinking and rigorous execution.
RESPONSIBILITIES:
-
Design, build, and operate petabyte-to-exabyte scale distributed systems for data acquisition, web crawling, preprocessing, filtering/classification, and multimodal pipelines (CPU/GPU workloads).
-
Architect high-performance search/retrieval engines (vector/hybrid/semantic) at trillion-document scale, integrating with LLMs/agents for truth-seeking, low-hallucination reasoning, and real-time knowledge access.
-
Develop reliable inference serving infrastructure: load balancing, autoscaling, KV cache, batching, fault-tolerance, monitoring (Prometheus/Grafana), CI/CD (Buildkite/ArgoCD), and benchmarking for 100% uptime and optimal tail latency.
-
Optimize low-level performance: CUDA kernels (GeMM, attention), Triton/CUTLASS extensions, quantization/distillation/speculative decoding, GPU memory hierarchy, and model-hardware co-design for next-gen architectures.
-
Innovate on compilers/runtimes (JAX/XLA/MLIR, custom features for Hopper/Blackwell), distributed profiling/debugging tools, and interconnect fabrics (copper/optical, 1.6T+, Ser Des/photonics, topology simulation, vendor roadmaps).
-
Manage complex workloads across clouds/clusters: orchestration (Kubernetes), data bookkeeping/verifiability, high-speed interconnect validation, failure analysis, and telemetry/automation for production reliability.
BASIC QUALIFICATIONS:
-
Strong systems engineering skills with proven impact on large-scale distributed infrastructure (data processing, search, inference, or cluster networking).
-
Proficiency in Python and at least one compiled language (Rust, C++, Go, Java); experience building bespoke libraries, optimizing performance, and debugging complex systems.
-
Hands-on experience with at least one key area: petabyte-scale data pipelines/crawling (Spark/Ray/Kubernetes), web-scale search/retrieval (vector DBs, ranking, RAG), inference optimization (SGLang, kernels, batching), compiler features (JAX/XLA), or high-speed interconnects (optical/copper, Ser Des, signal integrity).job
-
Deep understanding of distributed systems challenges: high-throughput ops/sec, latency/throughput tradeoffs, fault-tolerance, monitoring, and scaling to production billions-of-users or 100k+ GPUs.
-
Passion for AI infrastructure: keeping up with SOTA techniques, first-principles problem-solving, meticulous organization/bookkeeping, and delivering rigorous, high-quality results.
PREFERRED SKILLS AND EXPERIENCE:
-
Experience with multimodal data (images/video/audio), epistemics/truth-seeking in retrieval, or agentic systems (long-horizon reasoning, feedback loops).
-
Low-level optimizations: CUDA kernel development (Tensor cores, attention), GPU profiling (Nsight), low-precision numerics, or interconnect pathfinding (LPO/LRO/CPO, photonics).
-
Production expertise in inference reliability (0% error target), CI/CD for ML, or cluster networking (topology, vendor collaboration, failure root-cause).
-
Track record owning end-to-end projects in hyperscale environments, with strong debugging, vendor management, or open-source contributions (e.g., SGLang).
COMPENSATION AND BENEFITS:
$180,000 - $440,000 USD
Base salary is just one part of our total rewards package at xAI, which also includes equity, comprehensive medical, vision, and dental coverage, access to a 401(k) retirement plan, short & long-term disability insurance, life insurance, and various other discounts and perks.
*xAI is an equal opportunity employer. For details on data processing, view our *Recruitment Privacy Notice.
Total Views
0
Apply Clicks
0
Mock Applicants
0
Scraps
0
Similar Jobs
About xAI

xAI
Series BX.AI Corp., doing business as xAI, is an American company working in the area of artificial intelligence (AI), social media and technology that is a wholly owned subsidiary of American aerospace company SpaceX.
201-500
Employees
Austin
Headquarters
$50B
Valuation
Reviews
4.1
25 reviews
Work Life Balance
3.9
Compensation
4.6
Culture
4.3
Career
4.3
Management
3.5
83%
Recommend to a Friend
Pros
Strong engineering culture with focus on code quality
Competitive compensation packages with equity
Flexible remote work options and good work-life balance
Cons
Organizational changes and restructuring can be disruptive
Internal politics in some teams
Fast-paced environment with tight deadlines
Salary Ranges
0 data points
Junior/L3
Junior/L3 · Technical Writer
0 reports
$89,690
total / year
Base
-
Stock
-
Bonus
-
$76,237
$103,144
Interview Experience
5 interviews
Difficulty
3.0
/ 5
Duration
14-28 weeks
Interview Process
1
Coding Assessment
2
Live coding round
3
Technical Interview
4
Systems Design
Common Questions
Coding challenges
Technical problem solving
Systems design
Algorithm implementation
News & Buzz
Elon Musk’s SpaceX and xAI Are Planning a Megamerger of Rockets and AI - The Wall Street Journal
Source: The Wall Street Journal
News
·
7w ago
SpaceX reportedly mulling Tesla merger or tie-up with Elon Musk’s xAI firm - The Guardian
Source: The Guardian
News
·
7w ago
SpaceX and xAI could be merging. Why Elon Musk is doing it—and what might happen next - Fast Company
Source: Fast Company
News
·
7w ago
Why Elon Musk would want to merge SpaceX with xAI or Tesla - Business Insider
Source: Business Insider
News
·
7w ago


