Vision Language Action

Diverse Teleoperation System

Vision-Language-Action(VLA) 기반 로봇 조작 연구를 위해 텔레오퍼레이션을 활용한 데이터 생성 과정을 진행해 보았다. 다양한 방식의 텔레오퍼레이션 시스템을 구성하여 로봇 조작 시연 데이터를 수집하고, VLA 학습에 활용할 수 있도록 데이터 정리 및 변환 과정을 포함한 데이터 생성 파이프라인을 구축하여 적용해 보았다.

Apple Vision Pro
Leader Follow - Piper-Franka
Vive Tracker - RB-Y1
Manus Glove - Aidin Hand
Leader Follower - RB-Y1
Vive Tracker + Manus Glove - RB-Y1 + Aidin Hand

Difficulty of collection high-quality demonstrations

Plug insertion failure case (Teleoperation)

Reverse Playback Motion

Conventional data collection methods tend to include noise in precise manipulation stages because these stages are difficult to control, and non-experts are slower and produce lower-quality data.

In contrast, our method produces a cleaner and more stable data distribution, allowing even non-experts to collect high-quality demonstrations more easily.

The training results with GR00T N1.5 also show that our method is more effective at obtaining high-quality demonstrations in a shorter amount of time.

Model Application

다양한 VLA 모델을 로봇 조작 환경에 적용해 보았다. 여러 모델을 활용하여 로봇 조작 태스크에 대한 Inference를 수행하고 로봇 시스템에서의 동작을 확인하였다.

Pi 0.5 - OpenPI
GR00T N1.5 - Nvidia
SmolVla - Hugging Face

Force-aware Learning

Contatct-rich Manipulation

실제 환경과 접촉이 발생할 때, 접촉력에 대한 이해가 없으면 환경을 부수거나 로봇이 부서질 수 있음

Physical AI가 사람 또는 환경과 안전하게 상호작용 하기 위해선 이미지 뿐만 아니라 힘, 촉각 정보를 통한 환경 인지가 필요함

Force/Torque 데이터를 활용하면 환경과 접촉하는 작업을 안전하고 정밀하게 수행할 수 있음

Unified Robot Control-Policy

Share

Table of Contents

Other Researches

Dexterous Hand

Multi-DOF Gripper with suction fingertip 기존의 multi-fingered suction gripper 는

Vision Language Action

Diverse Teleoperation System Vision-Language-Action(VLA) 기반 로봇 조작 연구를 위해 텔레오퍼레이션을

Perception

6D Pose Estimation with Miniature 미니어처를 이용한 고중량물의 6D Pose

Go to Top