Hello, I'm Jeongeun Park

I am a Postdoctoral Researcher at Georgia Institute of Technology, working with Sehoon Ha. I received my Ph.D. from Korea University, where I worked with Sungjoon Choi.

I’m passionate about how robots can assist people in everyday life by collaborating and interacting naturally. In particular, I aim to investigate how foundation models can enable robots to understand, anticipate, and respond to human needs, making day-to-day tasks smoother and more intuitive.

News

  • [October 2026]: I joined Georgia Institute of Technology as a Postdoctoral Researcher, working with Sehoon Ha.
  • [September 2026]: SPARK, a simple post-training method for adapting pretrained visual representations to robot control, got accepted to CoRL 2026.
  • [August 2026]: I gave a talk titled "Towards Safe and Efficient Robot Foundation Models" at Sungkyunkwan University (SKKU) [slides].
  • [June 2026]: Two papers (GRIT on dexterous grasping and VIUGIC on transparent-object pose estimation) got accepted to IROS 2026.
  • [April 2026]: My open-source tutorials on VLA (LeRobot-MuJoCo Tutorial, LeRobot-MuJoCo-VLA Tutorial) gained significant community interest with 446 stars in total on GitHub!
  • [March 2026]: I gave a talk about robot foundation models at LG CNS.
  • [January 2026]: Our paper about design optimization got accepted to ICRA 2026.
  • [December 2025]: I have successfully defended my Ph.D. dissertation, “Towards Multimodal Foundation Models for Human-Centered Robot Behavior” at Korea University.
▼ Show older news

Experience

Publications

Learning Transferable World-Action Models from Task-Paired Human Videos

Learning Transferable World-Action Models from Task-Paired Human Videos

Preprint, 2026

WATCH pre-trains a world-action model on task-paired human videos with cross-video prediction: given one demonstration as reference, the model predicts the actions and future frames of another demonstration of the same task, learning interactions that transfer across scenes and demonstrators and boosting downstream robot imitation learning.

Retrieve, Don't Retrain: Extending Vision-Language-Action Models to New Tasks at Test Time

Retrieve, Don't Retrain: Extending Vision-Language-Action Models to New Tasks at Test Time

NeurIPS 2026 W NeurIPS 2026 Workshop on Robotics World Modeling

ReCAP is a retrieval-augmented policy that adapts vision-language-action models to new tasks at deployment without retraining: the frozen policy conditions on retrieved trajectories at every control step, so new tasks are absorbed by indexing data rather than updating parameters.

SPARK: Simple Post-training for Adapting pRetrained Knowledge to Robot Control

SPARK: Simple Post-training for Adapting pRetrained Knowledge to Robot Control

CoRL 2026 Conference on Robot Learning (CoRL), 2026

A simple post-training method that adapts pretrained visual foundation models for robot control by combining dynamics-aware abstraction with knowledge preservation, yielding compact yet semantically meaningful state representations. SPARK consistently improves success rates and generalization across robotics benchmarks and transfers to real-world manipulation.

Learning Dexterous Grasping from Sparse Taxonomy Guidance

Learning Dexterous Grasping from Sparse Taxonomy Guidance

Juhan Park, Taerim Yoon, Seungmin Kim, Joonggil Kim, Wontae Ye, Jeongeun Park, Yoonbyung Chai, Geonwoo Cho, Geunwoo Cho, Dohyeong Kim, Kyungjae Lee, Yongjae Kim, Sungjoon Choi
IROS 2026 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 2026

GRIT is a two-stage framework that learns dexterous grasping from sparse taxonomy guidance: it first predicts a taxonomy-based grasp specification from the scene and task context, and then generates continuous finger motions that preserve the intended grasp structure, enabling controllable grasp adjustment via high-level taxonomy selection.

Pose Estimation of Transparent Objects via Depth Completion and Confidence-Guided Registration

Pose Estimation of Transparent Objects via Depth Completion and Confidence-Guided Registration

Jeongeun Park*, Yeoncheol Jang*, Changin Kim, Youngjoon Yoo, Sungjoon Choi * Equal contribution
IROS 2026 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 2026

VIUGIC is a framework for transparent-object pose estimation that combines depth completion with uncertainty-aware registration, using confidence scores to down-weight unreliable depth predictions near refractive boundaries.

LEGO: Latent‑space Exploration for Geometry‑aware Optimization of Humanoid Kinematic Design

LEGO: Latent‑space Exploration for Geometry‑aware Optimization of Humanoid Kinematic Design

ICRA 2026 International Conference of Robotics and Automation (ICRA), 2026

Using screw-theory-based joint axis representation and isometric manifold learning, we construct a compact, geometry-preserving latent space of robot designs in which optimization is tractable. We then solve design optimization in this latent space using gradient-free optimization.

Towards Multimodal Foundation Model for Human-Centered Robot Behavior

Towards Multimodal Foundation Model for Human-Centered Robot Behavior

Jeongeun Park
Ph.D. Thesis, Korea University, 2026
Hierarchical Vision Language Action Model Using Success and Failure Demonstrations

Hierarchical Vision Language Action Model Using Success and Failure Demonstrations

CoRL 2025 W CoRL 2025 2nd Workshop on Safe and Robust Robot Learning for Operation in the Real World

We propose VINE, a dual-system framework that injects failure-aware reasoning into VLAs. System 1 executes grounded action chunks, while System 2 builds a tree of thought states and scores candidate subgoals using both success and failure data.

Self-supervised Visual State Representation Learning for robotics from Dynamic Scenes

Token Bottleneck: One Token to Remember Dynamics

NeurIPS 2025 The Thirty-Ninth Annual Conference on Neural Information Processing Systems (NeurIPS), 2025
ICLR 2025 7th Robot Learning Workshop: Towards Robots with Human-Level Abilities

ToBo uses extreme masked autoencoding to learn compact, temporally aware visual state embeddings that boost robotic manipulation and locomotion performance in both simulation and real-world.

A Unified Framework for Motion Reasoning and Generation in Human Interaction

A Unified Framework for Motion Reasoning and Generation in Human Interaction

ICCV 2025International Conference on Computer Vision (ICCV), 2025

MoLaM (Interactive Motion language model) integrates language and motion modalities to understand, generate, and control interactive motions in multi-turn conversations, addressing the scarcity of multi-turn interactive motion data with a synthetic dataset, INTER-MT2.

Towards Embedding Dynamic Personas in Interactive Robots: Masquerading Animated Social Kinematics (MASK)

Towards Embedding Dynamic Personas in Interactive Robots: Masquerading Animated Social Kinematics (MASK)

RA-L 2024ICRA 2025IEEE Robotics and Automation Letters (RA-L), 2024
International Conference of Robotics and Automation (ICRA), 2025
ICRA 2025 The 2nd Workshop on Nonverbal Cues for Human-Robot Cooperative Intelligence

Employing a persona-driven interactive framework to animate an anthropomorphic robotic system, enhancing audience engagement through non-verbal interactions

SPOTS: Stable Placement of Objects with Reasoning in Semi-Autonomous Teleoperation Systems

SPOTS: Stable Placement of Objects with Reasoning in Semi-Autonomous Teleoperation Systems

ICRA 2024International Conference of Robotics and Automation (ICRA), 2024

Introducing a teleoperation framework that enhances the 'place' task in robotics by combining simulation-driven stability verification with semantic reasoning from large language models.

CLARA: Classifying and Disambiguating User Commands for Reliable Interactive Robotic Agents

CLARA: Classifying and Disambiguating User Commands for Reliable Interactive Robotic Agents

RA-L 2024ICRA 2024IEEE Robotics and Automation Letters (RA-L), 2024,
International Conference of Robotics and Automation (ICRA), 2024

Introducing a method for interactive robotic agents using large language models (LLMs) to classify user commands as clear, ambiguous, or infeasible, enhancing reliability by leveraging uncertainty estimation, situational awareness, and user interaction for disambiguation.

SOCRATES: Text-based Human Search and Approach using a Robot Dog

Text-based Human Search and Approach using a Robot Dog

RO-MAN 2024International Conference on Robot and Human Interactive Communication (RO-MAN), 2024.

A robotic system using textual descriptions for human search and approach, combining language models for identifying targets and a hybrid learning framework for generating human-friendly robotic motions

Zero-shot Active Visual Search (ZAVIS): Intelligent Object Search for Robotic Assistants

Zero-shot Active Visual Search (ZAVIS): Intelligent Object Search for Robotic Assistants

ICRA 2023International Conference of Robotics and Automation (ICRA), 2023

A mobile robot system that uses free-form text for target object search using commonsene knowledge.

Elucidating Robust Learning with Uncertainty-Aware Corruption Pattern Estimation

Elucidating Robust Learning with Uncertainty-Aware Corruption Pattern Estimation

Pattern Recognition 2023Pattern Recognition, 2023

Uncertainty-aware robust learning

Towards Defensive Autonomous Driving: Collecting and Probing Driving Demonstrations of Mixed Qualities

Towards Defensive Autonomous Driving: Collecting and Probing Driving Demonstrations of Mixed Qualities

IROS 2022The 2022 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 2022

R3 Driving Dataset, a novel collection of driving data categorizing abnormal behaviors to enhance OOD detection for autonomous driving.

Learning Robot Structure and Motion Embeddings using Graph Neural Networks

Learning Robot Structure and Motion Embeddings using Graph Neural Networks

ICML 2022 WICML 2022 Workshop on Machine Learning for Computational Design (ICML'22-ML4CompDesign), 2022

Using graph neural networks (GNN) to find compact, low-dimensional embeddings of a robot’s kinematic structure and pose data.

Semi-Autonomous Teleoperation via Learning Non-Prehensile Manipulation Skills

Semi-Autonomous Teleoperation via Learning Non-Prehensile Manipulation Skills

ICRA 2022The 2022 IEEE Conference of Robotics and Automation (ICRA), 2022

Semi-Autonomous Teleoperation framework for non-prehensile manipulation tasks.


Reviewing activities