NeRF-4Scenes: A Video Dataset for Subjective Assessment of NeRF
Abstract
The dataset contains 36 NeRF-generated videos captured from four different indoor and outdoor environments: S1 for outdoor, S2 for auditorium, S3 for classroom, and S4 for lounge entrance. Each scene is trained using three NeRF models: Nerfacto as M1, Instant-NGP as M2, and Volinga as M3. Finally, each trained scene is rendered on three customized trajectories referred to as P1, P2, and P3. There are a total of 36 videos (4 scenes × 3 models × 3 paths) each having its own individual name. For example, video S1M1P1 corresponds to the outdoor scene (S1), which is trained on the Nerfacto model (M1), and rendered on the first camera path (P1). The dataset is available on DataverseNO
Full text
NeRF-4Scenes: A Video Dataset for Subjective Assessment of NeRF Shaira Tabassum, Seyed Ali Amirshahi Norwegian University of Science and Technology, Gjøvik, Norway Email: [email protected], [email protected] Data Collection and Preparation: The NeRF-4Scenes dataset is collected capturing four real-world scenes (one outdoor and three indoor environments) from the NTNU Campus in Gjøvik, Norway. The outdoor scene is captured using a 4K camera-equipped drone and indoor scenes are captured using a Nikon D850 camera. Each of the four scenes is captured in approximately 6-8 minutes videos. The captured video files are further processed, trained, and prepared in the following steps: ●In order to create a sparse set of input images, every 12th frame is extracted from the captured video files to generate image sequences. ●The camera poses for each image frame are then calculated through a Structure-from-Motion (SfM) tool - COLMAP [1]. ●The processed dataset are then fed to Nerfstudio API [2] where they are further trained on three NeRF models and generated three dynamic trajectory videos for each case. 1
The provided NeRF-4Scenes dataset includes all these generated video files trained on three NeRF models on three dynamic trajectories - producing nine videos from each scene. The complete training dataset: image files and camera poses, will be made available with a subsequent publication. Data Description: The dataset contains 36 NeRF-generated videos captured from four different indoor and outdoor environments: S1 for outdoor, S2 for auditorium, S3 for classroom, and S4 for lounge entrance. Each scene is trained using three NeRF models: Nerfacto as M1 [2], Instant-NGP as M2 [3], and Volinga as M3 [4]. Finally, each trained scene is rendered on three customized trajectories referred to as P1, P2, and P3. The dataset contains a total of 36 videos (4 scenes × 3 models × 3 paths) each having its own individual name. For example, video S1M1P1 corresponds to the outdoor scene (S1), which is trained on the Nerfacto model (M1), and rendered on the first camera path (P1). Figure.1 illustrates an overview of the NeRF-4Scenes dataset demonstrating the experiment design and naming convention of the video files. Figure.1: An overview of the NeRF-4Scenes dataset. The top row shows sample frames from each scene, while the lower blocks show how each scene is trained on three models and three paths, depicting the relationship between scenes, models, and paths. Each video is thus uniquely identified by its corresponding configuration, combining scene (S1-S4), model (M1-M3), and path (P1-P3). Data Organization: We have provided the video files in two folders: 1. Video_original: This folder contains all the videos in their original dimension (1920 × 1080) without further modification or resizing. 2. Video_resized: This folder contains the resized videos that we created matching the target monitor we used in our subjective experiment, i.e. 1900 × 1070 in our case. The observers were shown two of these resized videos side-by-side and asked to pick one. The subjective data collected from 18 observers altogether is provided in the file Subjective_Data.csv. An additional file Subjective_Data.xlsx is also included where observer's responses are stored in individual sheets, each containing six columns stating the following information: 1. Observer: Unique identifier for the observer. 2. Left Video: Video ID displayed on the left side of the screen during the experiment. 3. Right Video: Video ID displayed on the right side of the screen during the experiment. 4. Observer Preference: The video chosen by the observer from the given pair as their preference. 5. Elapsed Time: Time spent by the observer on each video pair. 6. Video Count: Number of times the observer replayed the video pair. 2
If you use our dataset, please cite: Tabassum, Shaira; Amirshahi, Seyed Ali, "Quality of NeRF Changes with the Viewing Path an Observer Takes: A Subjective Quality Assessment of Real-time NeRF Model", 2024 16th International Conference on Quality of Multimedia Experience (QoMEX) Tabassum, Shaira; Amirshahi, Seyed Ali, 2024, "NeRF-4Scenes: A Video Dataset for Subjective Assessment of NeRF", https://doi.org/10.18710/LFHFJN, DataverseNO, V1 Reference: [1] Schonberger, J. L., & Frahm, J. M. (2016). Structure-from-motion revisited. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 4104-4113). [2] Tancik, M., Weber, E., Ng, E., Li, R., Yi, B., Wang, T., ... & Kanazawa, A. (2023, July). Nerfstudio: A modular framework for neural radiance field development. In ACM SIGGRAPH 2023 Conference Proceedings (pp. 1-12). [3] Li, S., Li, C., Zhu, W., Yu, B., Zhao, Y., Wan, C., ... & Lin, Y. (2023, June). Instant-3d: Instant neural radiance field training towards on-device ar/vr 3d reconstruction. In Proceedings of the 50th Annual International Symposium on Computer Architecture (pp. 1-13). [4] Kang, X., Liu, K., Duan, J., Gong, Y., & Qiu, G. (2023, October). P2I-NET: Mapping Camera Pose to Image via Adversarial Learning for New View Synthesis in Real Indoor Environments. In Proceedings of the 31st ACM International Conference on Multimedia (pp. 2635-2643). 3