Meta
Meta

Research Scientist, Multi-Modal Human Understanding

RoleMachine Learning
LevelMid Level
LocationPittsburgh, Panama, United States
WorkOn-site
TypeFull-time
PostedToday
Apply now

About the role

Meta is seeking a Research Scientist to advance multi-modal AI technologies for human understanding and synthesis. In this role, you will develop Vision-Language Models (VLMs) and video foundation models that enable machines to perceive, interpret, and generate rich representations of human behavior, expression, and interaction. Your research will span multi-modal reasoning, video understanding, and generative synthesis, enabling more natural and intuitive human-computer interaction at scale.
Research Scientist, Multi-Modal Human Understanding Responsibilities
Design and implement novel multi-modal architectures that fuse vision, language, and temporal signals for holistic human understanding
Develop and train Vision-Language Models (VLMs) for tasks including visual question answering, image-text reasoning, and grounded human-centric understanding
Build video foundation models capable of temporal reasoning, action synthesis and long-form video synthesis with applications to human behavior synthesis
Research generative synthesis techniques for human-centric content including video generation, motion synthesis, and multi-modal content creation
Conduct rigorous experiments to evaluate model performance across diverse benchmarks, analyze failure modes, and iterate on architectures to improve accuracy and generalization
Contribute to the full research lifecycle from problem formulation and dataset curation through model development and evaluation

Minimum Qualifications:

Currently has, or is in the process of obtaining a Bachelor's degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience. Degree must be completed prior to joining Meta
2+ years of experience in multi-modal AI research, including hands-on work with Vision-Language Models, video understanding, or human-centric AI systems
2+ years of experience implementing and training large-scale neural networks using frameworks such as Py Torch, with experience on transformer-based architectures
Experience designing and executing experiments to evaluate multi-modal model performance, including quantitative analysis across vision, language, and video benchmarks
Experience writing production-quality or research-quality code in Python for multi-modal AI applications

Preferred Qualifications:

Experience developing or fine-tuning Vision-Language Models for human understanding tasks
Experience with video foundation models, temporal transformers, or large-scale video pretraining
Track record of contributing to published multi-modal AI research at venues such as CVPR, ICCV, or NeurIPS
Experience with generative models for human synthesis including diffusion models, GANs, or autoregressive models for video or motion generation

About Meta

Meta builds technologies that help people connect, find communities, and grow businesses. When Facebook launched in 2004, it changed the way people connect. Apps like Messenger, Instagram and Whats App further empowered billions around the world. Now, Meta is moving beyond 2D screens toward immersive experiences like augmented and virtual reality to help build the next evolution in social technology. People who choose to build their careers by building with us at Meta help shape a future that will take us beyond what digital connection makes possible today—beyond the constraints of screens, the limits of distance, and even the rules of physics.
For those who live in or expect to work from California if hired for this position, please click here for additional information.
United States of America: $122,000/year to $181,000/year + bonus + equity + benefits
Individual compensation is determined by skills, qualifications, experience, and location. Compensation details listed in this posting reflect the base hourly rate, monthly rate, or annual salary only, and do not include bonus, equity or sales incentives, if applicable. In addition to base compensation, Meta offers benefits. Learn more about benefits at Meta.

Equal Employment Opportunity:

Meta is proud to be an Equal Employment Opportunity employer. We do not discriminate based upon race, religion, color, national origin, sex (including pregnancy, childbirth, reproductive health decisions, or related medical conditions), sexual orientation, gender identity, gender expression, age, status as a protected veteran, status as an individual with a disability, genetic information, political views or activity, or other applicable legally protected characteristics. You may view our Equal Employment Opportunity notice here.
Meta is committed to providing reasonable accommodations for qualified individuals with disabilities and disabled veterans in our job application procedures. If you need assistance or an accommodation due to a disability, fill out the Accommodations request form.
Apply for this job
Take the first step toward a rewarding career at Meta.
Recruiters can view your conversations with AI. Using AI is optional and does not impact the outcome of your application process.Learn more about AI usage and settings.
Sign up
to ask follow-up questions

Required skills

Machine learning

Model evaluation

Data workflows

About Meta

Pittsburgh

Headquarters