Dataset Viewer
Auto-converted to Parquet Duplicate
Search is not available for this dataset
video
video
label
class label
1observation.images.head_left
1observation.images.head_left
1observation.images.head_left
1observation.images.head_left
1observation.images.head_left
1observation.images.head_left
1observation.images.head_left
1observation.images.head_left
1observation.images.head_left
1observation.images.head_left
1observation.images.head_left
1observation.images.head_left
1observation.images.head_left
1observation.images.head_left
1observation.images.head_left
1observation.images.head_left
1observation.images.head_left
1observation.images.head_left
1observation.images.head_left
1observation.images.head_left
1observation.images.head_left
1observation.images.head_left
1observation.images.head_left
1observation.images.head_left
1observation.images.head_left
1observation.images.head_left
1observation.images.head_left
1observation.images.head_left
1observation.images.head_left
1observation.images.head_left
1observation.images.head_left
1observation.images.head_left
1observation.images.head_left
1observation.images.head_left
1observation.images.head_left
1observation.images.head_left
1observation.images.head_left
1observation.images.head_left
1observation.images.head_left
1observation.images.head_left
1observation.images.head_left
1observation.images.head_left
1observation.images.head_left
1observation.images.head_left
1observation.images.head_left
1observation.images.head_left
1observation.images.head_left
1observation.images.head_left
1observation.images.head_left
1observation.images.head_left
1observation.images.head_left
1observation.images.head_left
1observation.images.head_left
1observation.images.head_left
1observation.images.head_left
1observation.images.head_left
1observation.images.head_left
1observation.images.head_left
1observation.images.head_left
1observation.images.head_left
1observation.images.head_left
1observation.images.head_left
1observation.images.head_left
1observation.images.head_left
1observation.images.head_left
1observation.images.head_left
1observation.images.head_left
1observation.images.head_left
1observation.images.head_left
1observation.images.head_left
1observation.images.head_left
1observation.images.head_left
1observation.images.head_left
1observation.images.head_left
1observation.images.head_left
1observation.images.head_left
1observation.images.head_left
1observation.images.head_left
1observation.images.head_left
1observation.images.head_left
1observation.images.head_left
1observation.images.head_left
1observation.images.head_left
1observation.images.head_left
1observation.images.head_left
1observation.images.head_left
1observation.images.head_left
1observation.images.head_left
1observation.images.head_left
1observation.images.head_left
1observation.images.head_left
1observation.images.head_left
1observation.images.head_left
1observation.images.head_left
1observation.images.head_left
1observation.images.head_left
1observation.images.head_left
1observation.images.head_left
1observation.images.head_left
1observation.images.head_left
End of preview. Expand in Data Studio

SharpaDex v1.0

A multimodal dataset of real-world teleoperated demonstrations for bimanual dexterous manipulation.

Website GitHub Dataset License

Overview · Get Started · Tasks · Examples · Features · Language · Citation

Overview of diverse manipulation tasks from head and wrist cameras

An animated overview of 56 tasks in a uniform 8 × 7 grid, with head-camera and wrist-camera observations distributed throughout the montage. It spans assembly, tool use, deformable-object manipulation, cleaning, material transfer, and long-horizon activities. Open MP4.

Overview

SharpaDex v1.0 is a real-world teleoperation dataset for bimanual dexterous manipulation. It contains 28,993 demonstration episodes across 59 manipulation tasks, comprising 32,396,172 frames (approximately 300.0 hours at 30 FPS).

Data were collected through human teleoperation of a real bimanual robotic system with two 7-DoF arms, two 22-DoF dexterous hands, four RGB cameras, and fingertip tactile sensors. Each episode provides synchronized joint state, joint torque, tool-center-point state, observed and commanded TCP poses, action, numeric tactile measurements, RGB video, tactile video, and temporally aligned language annotations.

The task set covers object rearrangement, articulated-object interaction, assembly, tool use, deformable-object manipulation, packaging, cleaning, and long-horizon sequential manipulation. The data are distributed in LeRobot v3.0 and v2.1 formats for research on imitation learning, visuomotor control, vision-language-action models, visual-tactile learning, and hierarchical policy learning.

Dataset at a Glance

Tasks Episodes Duration Frequency Video streams Language segments
59 28,993 300.0 hours 30 FPS 6 208,264
Modality Content
Robot 65D joint state, 65D action, 65D joint torque, 24D TCP state, and separate 12D observed / commanded TCP poses
Vision Two head cameras and two wrist cameras
Tactile 60D force/torque signal, deformation video, and raw tactile video
Language Structured task descriptions, frame-aligned subtasks, and 45 skill labels
Format LeRobot v3.0 and v2.1 for every released season

Version note: v3.0 and v2.1 are two exports of the same demonstrations. Dataset scale must be reported as 28,993 episodes and 32,396,172 frames, not the sum of both exports.

Detailed dataset statistics
Item Value
Manipulation tasks 59
Collection seasons 495
Episodes (one export) 28,993
Frames (one export) 32,396,172
Approximate duration at 30 FPS 300.0 hours
LeRobot v3.0 seasons / episodes 495 / 28,993
LeRobot v2.1 seasons / episodes 495 / 28,993
FPS 30
Synchronized video streams 6
State / action dimension 65
Joint-torque dimension 65
TCP-state dimension 24
Observed / commanded TCP-pose dimension 12 / 12
Tactile-signal dimension 60
Episodes with global language annotations 28,993 (100%)
Temporally grounded language segments 208,264
Frames covered by temporal language segments 28,897,608 (89.2%)
Distinct skill labels 45

For new projects, we recommend starting with lerobot_v3.0. The lerobot_v2.1 export is provided for compatibility with pipelines that depend on the older LeRobot layout.

Example Observations

The previews below show synchronized observations from a real-robot teleoperation episode of deal_playing_cards (season POC22007_2026_03_26_15_55_47_train, episode 12, source interval 5–25 seconds). Each clip is displayed at 2x playback speed. The MP4 links remain available when animated GIF playback is disabled by the Markdown viewer.

Head camera (MP4) Wrist camera (MP4)
Head-left camera observation Right-wrist camera observation
Tactile deformation (MP4) Raw tactile observation (MP4)
Tactile deformation observation Raw tactile observation

Task Collection

Task directories use lowercase snake_case names with an action-object-target structure. Counts refer to one format version (v3.0); v2.1 contains the corresponding demonstrations.

Complete task inventory — 59 tasks across 495 collection seasons
Task Seasons Episodes Frames
insert_batteries_into_charger 3 458 841,645
scrub_cup_interior 2 481 314,401
insert_plug_into_socket 6 416 183,726
assemble_gears_on_base 11 1,488 983,123
clean_plate_with_eraser 5 343 190,467
tighten_bottle_cap 7 131 139,693
secure_coiled_cable_with_velcro_tie 10 473 632,425
collect_waste_into_bin 8 508 621,140
deal_playing_cards 23 1,088 2,074,644
empty_dustpan_into_bin 10 488 379,585
cover_ball_with_cup 5 488 188,956
insert_knife_into_cutlery_bin 4 495 232,113
place_ball_in_box 3 115 40,672
place_ball_in_cup_then_cup_in_box 4 456 252,933
place_cup_on_plate 4 429 149,882
place_knife_and_fork_on_plate 8 461 314,324
stack_two_plates 4 491 185,481
remove_block_from_box 5 388 142,321
reorient_cup_upright 5 420 151,573
fit_trash_bag_into_bin 47 656 852,443
turn_book_pages 7 65 87,056
fold_and_close_box 7 537 1,008,242
hang_garment_on_hanger 15 521 979,299
form_kraft_paper_box 5 455 775,252
close_kraft_paper_box 5 645 693,683
fold_towel 3 459 547,201
grind_medicine_with_mortar_and_pestle 4 131 176,595
hammer_nails 4 1,071 861,668
insert_batteries_into_device 7 516 456,513
iron_garment 33 517 2,277,751
clean_garment_with_lint_roller 3 810 737,849
stack_napkins_in_holder 5 497 545,816
nest_cups 10 464 763,601
organize_objects_in_drawer 5 95 115,637
pack_blueberries_in_bag 17 485 1,533,417
pack_object_in_storage_bag 9 311 644,006
staple_paper 9 695 605,253
relocate_tennis_ball 1 450 157,112
place_fruits_in_basket 8 416 286,988
place_pens_in_holder 11 399 427,274
place_coffee_filter_in_dripper 21 784 1,786,030
insert_socket_module_into_board 2 957 512,619
transfer_pills_between_cups 9 479 239,561
pry_nails_from_board 3 24 21,174
reorient_and_relocate_carton 2 347 181,309
transfer_liquid_with_dropper 12 360 655,829
transfer_salt_between_cups_with_spoon 4 1,483 831,390
scoop_salt_from_jar_into_cup 13 424 485,424
sort_utensils 5 419 460,851
reposition_and_stack_blocks 15 1,007 1,069,334
sweep_waste_into_dustpan 10 609 516,260
seal_box_with_tape 1 115 208,648
place_books_on_shelf 9 151 133,230
collect_toys_into_basket 5 84 86,296
rotate_bottle_cap 17 583 497,695
unscrew_bottle_cap 3 445 351,637
clean_knife_with_cloth 2 292 246,287
play_tic_tac_toe 16 521 1,317,513
tomato_boxed_lunch_packaging 4 97 243,325

Get Started

Download the Dataset

Install Git LFS before cloning from Hugging Face.

git lfs install
git clone https://huggingface.co/datasets/Sharpa-Robotics/SharpaDex-v1.0

To clone metadata first and fetch large files later:

GIT_LFS_SKIP_SMUDGE=1 git clone https://huggingface.co/datasets/Sharpa-Robotics/SharpaDex-v1.0

For a single task, use sparse checkout. This example downloads clean_plate_with_eraser:

git init SharpaDex-v1.0
cd SharpaDex-v1.0
git remote add origin https://huggingface.co/datasets/Sharpa-Robotics/SharpaDex-v1.0
git sparse-checkout init
git sparse-checkout set clean_plate_with_eraser README.md
git pull origin main

Quick Inspection

Inspect meta/info.json to discover the exact schema and path templates for an export.

import json
from pathlib import Path

dataset_root = Path("SharpaDex-v1.0")
episode_root = (
    dataset_root
    / "clean_plate_with_eraser"
    / "season_POC22027_2026_04_17_11_01_59_train"
    / "lerobot_v3.0"
)

with open(episode_root / "meta" / "info.json", "r") as f:
    info = json.load(f)

print("episodes:", info["total_episodes"])
print("frames:", info["total_frames"])
print("fps:", info["fps"])
print("features:", list(info["features"]))

Dataset Structure

The repository uses a uniform task / season / format hierarchy. Every task is stored directly under the repository root.

Repository layout
SharpaDex-v1.0/
├── README.md
├── clean_plate_with_eraser/
│   ├── season_POC22027_2026_04_17_11_01_59_train/
│   │   ├── lerobot_v3.0/
│   │   │   ├── meta/
│   │   │   │   ├── info.json
│   │   │   │   ├── modality.json
│   │   │   │   ├── episodes/
│   │   │   │   ├── tasks.parquet
│   │   │   │   └── subtasks.parquet
│   │   │   ├── data/
│   │   │   │   └── chunk-000/
│   │   │   └── videos/
│   │   │       ├── observation.images.head_left/
│   │   │       ├── observation.images.head_right/
│   │   │       ├── observation.images.wrist_left/
│   │   │       ├── observation.images.wrist_right/
│   │   │       ├── observation.images.tactile_deform/
│   │   │       └── observation.images.tactile_raw/
│   │   └── lerobot_v2.1/
│   │       ├── meta/
│   │       ├── data/
│   │       └── videos/
│   └── season_.../
├── cover_ball_with_cup/
├── reposition_and_stack_blocks/
├── play_tic_tac_toe/
└── .../

Storage Layout

Part Description
meta/ Dataset schema, episode metadata, statistics, tasks, subtasks, and annotations
data/ Frame-aligned robot data stored as Apache Parquet files
videos/ Per-camera MP4 video streams
LeRobot path templates

LeRobot v3.0 uses paths similar to:

data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet
videos/{video_key}/chunk-{chunk_index:03d}/file-{file_index:03d}.mp4

LeRobot v2.1 uses paths similar to:

data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet
videos/chunk-{episode_chunk:03d}/{video_key}/episode_{episode_index:06d}.mp4

Feature Schema

All inspected seasons expose the same main feature set.

Feature Type Shape Description
observation.state float32 65 Joint-space robot state
action float32 65 Joint-space action target
observation.state.joint_torque float32 65 Joint-torque signal
observation.state.tcp float32 24 Left/right tool-center-point pose and force state
observation.state.tcp_pose float32 12 Left/right observed TCP poses, 6 components per arm
action.tcp_pose float32 12 Left/right commanded TCP poses, 6 components per arm
observation.tactile float32 60 10 fingertips x 6-axis force/torque signal
observation.images.* video varies Six synchronized visual and tactile streams
timestamp float32 1 Frame timestamp
frame_index int64 1 Frame index within an episode
episode_index int64 1 Episode index
task_index int64 1 Task-description index
subtask_index int64 1 Subtask-annotation index

Joint-Space State and Action

The 65D state, joint-torque, and action vectors use the following order:

Range Names Meaning
0-6 left_arm_j0 to left_arm_j6 Left arm joints
7-28 left_hand_j0 to left_hand_j21 Left dexterous-hand joints
29-35 right_arm_j0 to right_arm_j6 Right arm joints
36-57 right_hand_j0 to right_hand_j21 Right dexterous-hand joints
58-64 motor_j0 to motor_j6 Torso / motor-related joints

TCP Pose State and Action

observation.state.tcp_pose and action.tcp_pose are separate 12D float32 vectors. Components 0:6 contain the left-arm TCP pose and 6:12 contain the right-arm TCP pose.

The observation is read from the source state/left_arm/tcp_pose and state/right_arm/tcp_pose; the action is read from the corresponding action/.../tcp_pose streams. Each stream uses its own aligned_index. Source values and pose representation are preserved; commanded poses are not copied from observed poses or reconstructed using forward kinematics.

The existing 24D observation.state.tcp remains unchanged: left pose 0:6, left force/torque 6:12, right pose 12:18, right force/torque 18:24. Its two pose slices correspond to the new observed TCP-pose vector.

The export does not fully declare TCP units, reference frames or rotation convention (rotation_type is null). Use the acquisition/controller convention before geometric transformations; do not infer Euler angles or quaternions from the vector width alone.

Video Streams

Feature key Description Shape
observation.images.head_left Left head camera 480 x 480 x 3
observation.images.head_right Right head camera 480 x 480 x 3
observation.images.wrist_left Left wrist camera 480 x 480 x 3
observation.images.wrist_right Right wrist camera 480 x 480 x 3
observation.images.tactile_deform Tactile deformation video 480 x 1200 x 3
observation.images.tactile_raw Raw tactile video 480 x 1600 x 3

Video files are MP4 without audio. Codec may differ between exports; inspect the local files if your training stack has codec restrictions.

Language Annotations

Every episode includes language annotations at two complementary levels: a global description of the complete task and temporally grounded descriptions of the subtasks performed during the trajectory. These annotations support language-conditioned policy learning, vision-language-action training, temporal grounding, skill discovery, hierarchical policy learning, and subtask-aware evaluation.

Global Task Description

The episode-level global_task stores the available structured natural-language fields below:

Component Meaning Example
Task Short task name ClearPlate
Instruction Ordered description of the intended behavior Pick up the plate and eraser, wipe the plate, then place both objects down
Scene Objects, layout, and manipulation context A marked plate and a blackboard eraser are on the table
Success Completion criteria The marks are removed and both objects are placed stably

The same information is also exposed as separate fields under tags. All episodes have scene descriptions; 1,061 scene descriptions were supplemented from existing episode text and visual spot checks, with provenance recorded in the annotation metadata. Some of these episodes do not have separate instruction, success-criteria or SOP fields:

task_name
task_instruction
scene_description
success_criteria
score
SOP

Temporally Grounded Subtasks

Each episode is decomposed into language segments. A segment contains:

Field Description
start_step, end_step Frame-index interval for the subtask
text Natural-language description of the action phase
task_skill Compact semantic skill label such as pick, place, move, insert, fold, or wipe
label Annotation type, including subtasks and recovered action descriptions
id Episode-local subtask identifier

Across the release, all 28,993 episodes contain language annotations. The dataset provides 208,264 temporal subtask segments spanning 28,897,608 frames, or approximately 89.2% of all frames. The segments use 45 nonempty skill labels; recovered descriptions may not have a separate skill label. Frames outside a fine-grained segment remain associated with the episode-level task description.

For example, a plate-cleaning trajectory is decomposed into phases such as:

Pick up the plate from the table with the left hand. | Skill: pick
Pick up the blackboard eraser with the right hand. | Skill: pick
Wipe away the dirty marks while the left hand holds the plate. | Skill: wipe
Place the plate and the eraser back on the table. | Skill: place

Annotation Files and Frame Alignment

Metadata Purpose
meta/annotations.jsonl Per-episode global task, structured tags, temporal segments, frame ranges, and source identifiers
meta/tasks.jsonl / tasks.parquet Mapping from task_index to the global task text
meta/subtasks.jsonl / subtasks.parquet Mapping from subtask_index to subtask text and skill
Frame feature task_index Connects each frame to its global task description
Frame feature subtask_index Connects each frame to its subtask annotation

The v2.1 export stores task and subtask lookup tables as JSONL; v3.0 stores the corresponding tables as Parquet. annotations.jsonl is included with both exports for direct episode-level inspection.

The following example reads the global description and temporally grounded language from the first annotated episode:

import json
from pathlib import Path

meta_root = Path(
    "SharpaDex-v1.0/clean_plate_with_eraser/"
    "season_POC22027_2026_04_17_11_01_59_train/"
    "lerobot_v3.0/meta"
)

with open(meta_root / "annotations.jsonl", "r") as f:
    annotation = json.loads(next(f))

print(annotation["global_task"])

for segment in annotation["language_segments"]:
    print(
        segment["start_step"],
        segment["end_step"],
        segment["task_skill"],
        segment["text"],
    )

When constructing training samples, use task_index for episode-level language conditioning and subtask_index or the explicit start_step / end_step intervals for phase-level conditioning.

Tactile Modality

The dataset provides tactile information in structured numeric and image-based forms:

  • observation.tactile contains 60 values: left/right thumb, index, middle, ring, and little fingertips, each with fx, fy, fz, tx, ty, and tz.
  • observation.images.tactile_deform visualizes contact-induced deformation patterns.
  • observation.images.tactile_raw preserves the raw tactile camera layout for custom preprocessing and representation learning.

All tactile modalities are synchronized with visual observations, proprioception, and actions at 30 FPS. The 60D signal provides a compact tactile representation; the tactile video streams provide spatially resolved observations for models designed to process the additional input resolution.

Usage Recommendations

For policy learning, a typical configuration is:

  • Visual observations: one or more observation.images.* streams
  • Proprioception: observation.state
  • Optional force and contact inputs: joint torque, TCP state, and observation.tactile
  • Optional tactile vision: observation.images.tactile_deform and/or observation.images.tactile_raw
  • Supervision target: action, or action.tcp_pose for a TCP-based action representation
  • Language conditioning: structured task instruction, scene description, and success criteria
  • Phase conditioning: frame-aligned subtask text and skill label

For evaluation, split by collection season rather than randomly splitting frames. This reduces temporal leakage and avoids placing closely related demonstrations from the same collection session in both training and evaluation sets. For multi-task experiments, additionally report held-out tasks or task families when measuring cross-task generalization.

Dataset Notes

  • The release is organized by task and collection season, not as one flattened LeRobot root.
  • Every released season includes both lerobot_v3.0 and lerobot_v2.1.
  • Task descriptions, scene details, success criteria, and subtask decomposition may vary between seasons; use the metadata shipped with the selected season as the source of truth.
  • Some long-horizon tasks contain multiple interaction phases and substantial hand-object occlusion.
  • The two tactile video streams are high resolution and may dominate input bandwidth.
  • The dataset contains demonstrations rather than a fixed benchmark split; users should document their task and season splits for reproducibility.

License and Terms

This dataset is released under the Creative Commons Attribution 4.0 International License (CC BY 4.0). You may share and adapt the dataset, including for commercial purposes, provided that you give appropriate attribution and indicate whether changes were made.

Citation

If this dataset contributes to your research, please cite or acknowledge the dataset:

@misc{sharpadex_v1_2026,
  title        = {SharpaDex v1.0: Large-Scale Bimanual Dexterous Manipulation Demonstrations},
  author       = {{Sharpa}},
  howpublished = {\url{https://huggingface.co/datasets/Sharpa-Robotics/SharpaDex-v1.0}},
  year         = {2026}
}
Downloads last month
48