Jaemo Jeong
gosfl4760@kaist.ac.kr
KAIST M.S. researcher developing multimodal perception under practical constraints, with work spanning training-free audio-visual event localization and alignment of exocentric video with ambient sensors.
Research interests include multimodal foundation models and world models for embodied perception in robotics and autonomous systems.
Research interests include multimodal foundation models and world models for embodied perception in robotics and autonomous systems.
Work Experience
KAIST
Daejeon, Republic of Korea
Researcher (M.S. Student)
Mar 2025 - now
- Research on multimodal perception under practical constraints, including:
- Training-free audio-visual event localization
- Alignment of exocentric video with ambient sensors
- Developing foundation models and world models for embodied perception in robotics and autonomous systems.
Education
KAIST
Daejeon, Republic of Korea
Computer Science (M.S.)
Mar 2025 - Feb 2027
M.S. in Computer Science (expected Feb 2027)
KAIST
Daejeon, Republic of Korea
Computer Science (B.S.)
Mar 2020 - Feb 2025
B.S. in Computer Science (Mar 2020 – Feb 2025)
Skills
Programming Languages
- Python
- C/C++
Deep Learning & Frameworks
- PyTorch
- CLIP
- CLAP
Sensors & Signal Processing
- IMU
- Ambient sensors
Optimization & Localization
- Inference optimization
- Temporal localization
- Sensor fusion
Awards
KAIST Alumni Academic Scholarship
Daejeon, Republic of Korea
2024 KAIST Alumni Academic Scholarship
Jan 2024
Projects
INT16 Systolic Array Accelerator with YOLO-tiny Inference
Hardware systems project
Sep 2023 - Dec 2023
- Implemented an INT16 FPGA systolic array accelerator for YOLO-tiny inference using BRAM tiling and both weight-stationary and output-stationary dataflows.
Dayscout, Kakao Impact Tech for Impact Campus
Backend Developer | Finalist
Sep 2023 - Dec 2023
- Built FastAPI services for a community-curated nutrition wiki for people with type 1 diabetes.
- Parsed CLOVA OCR output into standardized nutrition records and implemented MySQL food search.
Publications
- [NeurIPS 2026 under review] Jaemo Jeong, Junho Yoon, Hyunju Kim, and Dongman Lee.
- Identified false co-activation, where semantically related absent labels can outrank weaker present labels and cannot be corrected by scalar thresholding.
- Developed two-stage training-free inference over frozen embeddings using nonnegative Lasso, reliability-weighted cross-modal priors, and modality-specific re-selection.
- Improved LLP Type@seg by 5.94 points and OV-AVEBench Avg by 9.32 points over same-backbone baselines, adding 4.27 ms per sample after shared encoding.
- [CVPR 2026] Junho Yoon*, Jaemo Jeong*, Hyunju Kim*, and Dongman Lee.
- Studied activity recognition using fixed exocentric video and object-mounted ambient sensors, removing the need for body-worn cameras and sensors.
- Learned decomposed spatial and temporal alignment for sensor-only, video-only, and joint inference after multimodal training. Sensor-only inference outperformed adapted baselines by up to 30% in F1 and 50% in mAP.
- In preparation (Targeting ICLR 2027)
- Developing causal streaming inference with frozen encoders and bounded event memory.
- B.S. thesis
- Reimplemented InfiniPot's CaP and NuC token selection methods without released code in a Hugging Face BERT prototype.
- Designed a Hit Ratio diagnostic showing that token overlap becomes a weaker proxy for semantic preservation on longer sequences.