JPMorgan Chase
JPMorgan Chase

Vice President - Lead Infrastructure Engineer | Storage

RoleInfrastructure
LevelLead
LocationHouston, United States
WorkOn-site
TypeFull-time
Posted2 months ago
Apply now

About the role

Assume a vital position as a key member of a high-performing team that delivers infrastructure and performance excellence. Your role will be instrumental in shaping the future at one of the world's largest and most influential companies.

As a Lead Infrastructure Engineer at JPMorgan Chase within the Enterprise Technology, Infrastructure Platforms team, you apply deep knowledge of software, applications, and technical processes within the infrastructure engineering discipline. Continue to evolve your technical and cross-functional knowledge outside of your aligned domain of expertise.

Job responsibilities

  • Uses enterprise-authorized AI capabilities within the work environment to accelerate infrastructure analysis and design documentation, validating outputs and handling operational data according to sensitivity and security requirements.
  • Applies reuse-first, AI-assisted practices within delivery and automation routines to identify recurring issues and validate remediation options, ensuring changes are traceable/auditable and aligned to resiliency and security expectations.
  • Own and continuously improve SLOs/SLIs, error budgets, on-call readiness, and operational excellence for storage services.
  • Lead incident response for storage outages/performance degradations; drive RCAs and implement preventative actions.
  • Create and maintain runbooks, escalation paths, and standardized operational procedures.
  • Operate and enhance block/file/object storage platforms across on-prem and/or cloud environments.
  • Perform performance tuning, capacity planning, lifecycle management, and resiliency testing (failover/DR validation).
  • Partner with infrastructure, network, OS, database, and application teams to meet workload requirements and reliability targets.
  • Build automation for provisioning, patching, upgrades, replication, backup/restore, and compliance checks.
  • Implement AI-driven observability/AIOps (telemetry correlation, anomaly/regression detection, LLM-assisted incident/runbook workflows) with accuracy, auditability, and safe rollout.

Required qualifications, capabilities, and skills

  • Formal training or certification on infrastructure engineering concepts and 5+ years applied experience
  • Demonstrated experience using enterprise-authorized AI capabilities within the work environment to support infrastructure engineering workflows with strong validation habits and awareness of data sensitivity.
  • Ability to review and validate AI-assisted recommendations before implementation, escalating when uncertain and ensuring outcomes align to resiliency, security, and auditability expectations.
  • Strong knowledge of storage fundamentals (RAID/erasure coding, replication, snapshots, tiering/caching, IOPS/latency, multipathing, SAN/NAS, object semantics).
  • Hands-on experience with at least one major storage ecosystem (e.g., Net App, Dell EMC Power Store/Isilon, Pure, Hitachi, Ceph, IBM, or cloud storage services).
  • Solid Linux fundamentals, including system performance, networking basics, and kernel/storage-stack concepts.
  • Strong scripting/programming in one or more of Python, Go, Bash.
  • Experience with observability stacks (e.g., Prometheus/Grafana, ELK/Open Search, Splunk, Datadog, OpenTelemetry).
  • Proven incident management skills and ability to operate effectively in an on-call rotation.
  • Practical AI/data skills for operations (anomaly detection/forecasting/correlation/classification; feature extraction and evaluation; integrating AI into production tooling/CI/CD; safe LLM use with guardrails and human-in-the-loop review).

Preferred qualifications, capabilities, and skills

  • Kubernetes storage (CSI), stateful workloads, and container platform operations.
  • Infrastructure as Code (Terraform/CloudFormation) and configuration management (Ansible/Chef/Puppet).
  • Streaming/queue tooling for telemetry and event pipelines (e.g., Kafka).
  • Experience with ITSM/event management platforms (e.g., Service Now).
  • Backup/DR products and strategy design, including RPO/RTO tradeoffs.
  • Security controls for data platforms (KMS/HSM, secrets management, key rotation).
  • Experience building/operating controlled self-service platforms with guardrails to reduce toil at scale.

Benefits and perks

Healthcare

Paid Time Off

Retirement Plan

Learning Budget

Equity

Required skills

Systems engineering

Testing

Technical planning

About JPMorgan Chase

Houston

Headquarters