HyperAIHyperAI

Command Palette

Search for a command to run...

SceneFun3D 3D Scene Function Interaction Dataset

Date

an hour ago

License

Non-Commercial

SceneFun3D is a 3D scene functional interaction dataset released by Voxel51 in 2024. It focuses on fine-grained understanding of small interactive parts (such as handles, knobs, and switches), covering part localization, Gibsonian functional reasoning, motion parameter estimation, and natural language task description. Related research papers are published under the title "SceneFun3D". Fine-Grained Functionality and Affordance Understanding in 3D Scenes . This dataset encompasses 710 high-resolution real-world indoor scenes, containing 14,867 functional interaction element annotations, 9 functional categories, 14,279 motion parameters, and over 10,913 natural language task instructions. Built upon high-resolution Faro laser scanning point clouds and iPad video sequences, the dataset is divided into a training set of 545 scenes, a validation set of 80 scenes, and a test set of 85 scenes. This version samples 10 scenes from each category as the benchmark test set; the functional annotations in the test set have been hidden.

Dataset composition:

  • Scene grouping: The dataset adopts a grouped dataset structure, with each group corresponding to an independent scene, identified by visit_id.
  • 3D point cloud data: Provides downsampled (5 mm voxel) FARO laser scan point cloud, including annotations of functional interactive elements, ARKit room-level object bounding boxes, and task descriptions.
  • Video sequence data: Each scene is accompanied by 2-3 high-resolution RGB video sequences (1920 x 1440 resolution, approximately 10 FPS) recorded on an iPad Pro, which also include frame-by-frame depth maps, camera poses and intrinsic parameters.

Data fields:

  • laser_scan slices: contain functional_elements (3D functional element detection boxes and attributes), objects_3d (ARKit object boxes), and tasks (natural language task list).
  • ipad_N slice: contains frame-by-frame depth (depth heatmap), intrinsics (camera intrinsics), camera_pose (camera pose), and projected_elements/projected_objects (projected bounding boxes and keypoints of 3D elements/objects in 2D frames).
  • Functional element attributes: Each 3D detection box contains a label (9 functional categories), motion_type (translation/rotation), motion_dir (motion direction), motion_origin (motion origin), and descriptions (task instructions).

Citation

@inproceedings{delitzas2024scenefun3d,
title={{SceneFun3D: Fine-Grained Functionality and Affordance Understanding in 3D Scenes}},
author={Delitzas, Alexandros and Takmaz, Ayca and Tombari, Federico and Sumner, Robert and Pollefeys, Marc and Engelmann, Francis},
booktitle={IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)},
year={2024}
}

Build AI with AI

From idea to launch — accelerate your AI development with free AI co-coding, out-of-the-box environment and best price of GPUs.

AI Co-coding
Ready-to-use GPUs
Best Pricing

HyperAI Newsletters

Subscribe to our latest updates
We will deliver the latest updates of the week to your inbox at nine o'clock every Monday morning
Powered by MailChimp