Documentation Index

Fetch the complete documentation index at: https://docs.dataloop.ai/llms.txt

Use this file to discover all available pages before exploring further.

Datasets Overview

Prev Next

Dataset

A dataset is a structured collection of files, along with their associated metadata and annotations. Datasets are organized in a folder hierarchy and can contain any number of nested subfolders, making it easy to manage and categorize data.

Supported File Types in Datasets

DDOE supports a broad range of file formats across multiple domains, enabling flexible and scalable data annotation workflows.

Learn here for details.

What Can You Do with a Dataset?

Datasets are the central workspace for managing, organizing, and preparing data for annotation and AI workflows. With a dataset, you can:

  • Create and Ingest Data: Upload files, synchronize cloud storage, connect on-premises storage, or integrate with Compute Cluster storage.

  • Manage Dataset Contents: Browse, search, filter, organize, move, clone, and delete files and folders.

  • Launch Annotation Workflows: Create annotation tasks, run pipelines, and trigger automated workflows using selected data.

  • Manage Metadata and Annotations: View, edit, import, and export metadata, annotations, and activity logs.

  • Analyze Data Quality: Detect duplicate files, missing annotations, unlabeled assets, and metadata inconsistencies.

  • Automate Preprocessing: Configure applications and pipelines to automatically process newly added data.

  • Explore Dataset Embeddings: Visualize and analyze embeddings to better understand data relationships and patterns.

  • Integrate Storage Systems: Connect datasets to AWS, Azure, Google Cloud, and on-premises storage platforms.

  • Monitor and Collaborate: Track dataset usage, monitor annotation progress, and share datasets with authorized users and teams.

  • Prepare AI Training Data: Organize and curate data for machine learning training, validation, and testing workflows.

Supported Data Storage Options for Datasets

Datasets can be stored in DDOE-managed internal storage or linked to external storage systems through storage integrations. External storage connections allow datasets to synchronize and manage data directly from cloud or on-premises storage without duplicating the source files. DDOE also supports advanced dataset operations such as cloning and merging, enabling efficient collaboration and data governance across both internal and external storage environments.

DDOE Internal Storage

Store dataset files directly within DDOE-managed storage. This option provides a fully managed storage environment where data is uploaded, stored, and managed within DDOE, simplifying dataset administration and reducing dependency on external storage infrastructure.

Cloud Storage Providers

Connect datasets to cloud-based storage services and manage data directly from the source location without moving or duplicating files.

AWS S3

Connect datasets with Amazon Simple Storage Service (S3) buckets. DDOE can access and synchronize files stored in AWS, enabling organizations to use existing cloud storage as a dataset source.

Azure Blob Storage

Connect datasets to Azure Blob Storage containers to access and manage data stored within Microsoft Azure environments while maintaining centralized dataset management in DDOE.

Google Cloud Storage (GCS)

Connect datasets to Google Cloud Storage buckets, allowing DDOE to access and manage data stored in Google Cloud Platform (GCP) without additional data migration.

Compute Cluster Storage

Connect datasets to storage resources available within a Compute Cluster. This allows DDOE to access and synchronize data that is already used for data processing, machine learning training, inference, or other compute-intensive workloads.

On-premises storage systems

Connect datasets to storage systems hosted within your organization's infrastructure. This enables secure access to enterprise data while keeping files within on-premises environments. The main types are NFS (Network File System), NFS with MetadataIQ (Dell PowerScale OneFS), S3-compatible API (Object Storage), and S3 API with MetadataIQ (Dell PowerScale OneFS).

Storage Type

Description

Key Benefits in DDOE

NFS (Network File System)

Connect to file shares using the NFS protocol. DDOE can access files directly from shared storage locations, making it easy to use existing enterprise file systems as dataset sources.

• Direct access to files without data migration
• Simple integration with existing enterprise storage
• Supports large file repositories and folder hierarchies
• Keeps data within the organization's network
• Enables centralized dataset management across shared storage

NFS with MetadataIQ (Dell PowerScale OneFS)

Combines direct file access through NFS with MetadataIQ's metadata indexing capabilities. DDOE can access files from PowerScale storage while using MetadataIQ to accelerate asset discovery, metadata retrieval, and dataset synchronization.

• Direct access to files through NFS
• Faster dataset discovery and synchronization
• Improved metadata-based search and filtering
• Reduces storage scanning overhead
• Accelerates dataset creation and refresh operations
• Scales efficiently for large image, video, document, and LiDAR datasets

S3-Compatible API (Object Storage)

Connect to any object storage platform that supports the Amazon S3 API. This enables DDOE to integrate with private cloud, enterprise, and third-party object storage solutions using a standard S3-compatible interface.

• Standard S3 interface for broad storage compatibility
• Access data without copying it into DDOE storage
• Supports scalable object-based storage architectures
• Simplifies integration with private and enterprise clouds
• Enables centralized management of distributed data assets

S3 API with MetadataIQ (Dell PowerScale OneFS)

Connect to Dell PowerScale OneFS object storage using its S3-compatible API while leveraging MetadataIQ for advanced metadata indexing and discovery.

• Scalable object storage access through the S3 API
• Rapid asset discovery using metadata indexing
• Faster dataset synchronization and refresh operations
• Metadata-driven search and filtering for large datasets
• Improved visibility and governance of stored assets
• Optimized for enterprise-scale data management and collaboration

MetadataIQ

MetadataIQ is a Dell PowerScale OneFS capability that indexes file and object metadata, making large-scale storage environments easier to search and manage. Rather than scanning the entire file system, DDOE can query MetadataIQ indexes to rapidly discover assets and retrieve metadata.

How MetadataIQ Is Used in DDOE

When connecting Dell PowerScale OneFS storage to DDOE, MetadataIQ helps DDOE efficiently discover and manage dataset assets stored in on-premises environments. By leveraging the metadata index provided by MetadataIQ, DDOE can:

  • Accelerate dataset discovery without performing full storage scans.

  • Improve search and filtering performance across large datasets.

  • Access file metadata at scale for data management and governance.

  • Identify and organize assets more efficiently during dataset creation.

  • Synchronize storage contents faster when importing or refreshing datasets.

  • Reduce storage access overhead by querying metadata instead of repeatedly traversing the file system.

This approach is particularly beneficial for organizations managing large volumes of image, video, LiDAR, and document data within Dell PowerScale storage environments.

Learn here to create datasets.


Datasets Details

In the tab, the Datasets are displayed in a list view. The following list provides the list of available fields and specific criteria of search and filters:

  • To search: You can search datasets by Dataset Name.

  • To Filter: You can filter the listed datasets by the following criteria:

    Type

    Provider

    Driver Type

    The type of the datasets, whether the dataset is cloned, merged, or the original (master).

    • Master

    • Clone

    • Merge

    The available storage providers for the datasets.

    • DDOE

    • AWS

    • GCP

    • Azure

    The type of driver used from the storage provider.

    • File System

    • S3 Bucket

    • GCS Bucket

    • Blob Storage

    • Data Lake Storage Gen2

  • Select Creators: It allows you to filter datasets based on the creator.

List of Fields

The column values are populated according to the datasets.

Column Name

Description

Provider

It displays the name of the storage provider.

Dataset Name

The name of the dataset. Clicking on it will open the Data Browser page.

Items

The number of items available in the dataset.

Feature Sets

The number of Feature Sets available in the dataset.

Annotated

It displays the percentage of items that are annotated.

Type

It displays the type of the dataset, whether it is master (original), cloned, or merged.

Driver Type

It displays the name of the storage driver type.

Open Tasks

It displays the number of the tasks that are open.

Created at

The creation date of the dataset.

Created by

The Avatar of the user who created the dataset. You can see the email ID of the user when you hover.

Clicking on a dataset will displays the following features of the dataset:

The Dataset page provides access to all Datasets in the project. Datasets are listed in a customizable table:

  • Show/hide standard columns according to fields used.

  • Add custom columns to better manage datasets.

Custom Dataset Fields

You can add your context to Datasets to manage them in your projects according to your needs. Context is added as user Meta-Data in the Dataset entity (any Meta-data field outside the System area). These fields can then be reflected as columns in the Datasets page, presenting the information and context, allowing for Datasets to be sorted and filtered by these fields.


Dataset Actions

DDOE allows you to perform the following actions on your datasets. When you click on the three dots icon, the following options are displayed.

  • Merge Datasets: It allows you to merge two or more datasets. Select two or more datasets, and click Merge Datasets.

  • Sync Now: It allows you to start syncing data from your external cloud storage to the current dataset. Any changes in your storage will be applied to the dataset. This field is displayed only for the datasets belong to external cloud storages.

  • Upload: It allows you to upload files and folders to the selected dataset.

  • Export: It allows you to export the datasets in a zip file.

  • Clone: It allows you to clone two or more datasets after entering necessary details on the Clone Datasets/Items window.

  • Copy ID: It allows you to copy the ID of the dataset.

  • Rename: It allows you to rename the dataset.

  • View Embeddings: The Embeddings in the DDOE platform provides a powerful way to visualize and interact with feature sets derived from your datasets.

  • Related-Tasks Analytics: Clicking on the icon allows you to open and view the Analytics page of the selected dataset. It gives an insight of the related tasks.

  • Recipe: Clicking on the Recipe allows you to open the Recipe page.

  • Delete All Annotations: It deletes all the annotations of the items inside the selected dataset.

  • Delete Dataset: It allows you to delete the selected dataset.


Dataset Types

Deriving from its data-versioning, there are different types of Datasets:

  • Master: Original dataset that manages the actual binaries.

  • Clone: Contains pointers to original files, enabling management of virtual items that do not replicate the binaries of the underlying storage once cloned or copied. When you clone a dataset, you can decide whether the new copy will contain metadata and annotations created over the original.

  • Merge: Multiple datasets can be merged into one, which enables multiple annotations to be merged onto the same item.

Binaries dataset

The Binaries' dataset visible in your DDOE project is a system-generated dataset designed for storing binary files associated with the project, such as model binaries. While this dataset is created automatically and is not intended for direct user interaction, it can be viewed through the SDK or API.


Create Datasets

DDOE allows you to create datasets on the DDOE platform based on the Storage requirement.


Lock Datasets During Export

Ensure that datasets being exported remain consistent by preventing changes to items or annotations during the export process, while minimizing disruption to other users. When exporting a dataset, selecting the Lock dataset during export option will move the dataset into Read-Only Mode. This prevents any new actions and halts ongoing activities (e.g., running pipelines) until the .zip file compression is complete.

Learn how to lock datasets during export via UI and SDK.

Read-Only Mode for Locked Datasets

Read-Only Mode prevents accidental edits when datasets are locked due to exports or system restrictions, ensuring data integrity and controlled access.

How It’s Indicated:

  • A banner notification appears across all annotation studios and the dataset browser.

  • Positioned above save and status buttons.

  • Can be temporarily hidden but reappears upon refresh.

User Permissions:

  • Annotators: See “Saving is currently unavailable” with a tooltip and refresh button.

  • Managers: Same as annotators, but with export initiator details.

  • Developers & Admins: Same as managers, plus an Unlock button to remove temporary locks.

Restricted Actions:

  • Save and status buttons are disabled.

  • Auto-save is turned off to avoid repeated error messages.

  • Pipeline executions that modify the dataset will fail during the lock.

  • Error messages displayed when users attempt modifications: Action failed: The dataset is locked during export. Please try again after the export is finished.

Unlocking (For Developers Only)

  • Click Unlock on the read-only banner. A confirmation dialog appears.

  • Click Confirm to release the lock. The dataset becomes editable.

Refreshing Dataset Status

  • Click Refresh to check the latest dataset status without a full page reload.

  • The system auto-checks every 30 seconds for updates.

Automatic lock timeout for stuck exports

The Automatic Lock Timeout prevents datasets from being indefinitely locked due to failed or stuck exports. The system automatically releases dataset locks after the timeout (2 hours), ensuring smooth operations.

  • A lock is placed on a dataset when an export starts.

  • If the export fails or gets stuck, the lock is automatically released after a set time in seconds based on dataset size (e.g., lock_timeout_sec = 1000).

  • Admins can adjust the timeout duration.

  • Use the SDK to set the time.

User notifications

  • Error Message: The previous export lock expired. Please try again.

  • System Log Warning: Export lock timeout exceeded for Dataset [ID]. Lock released.

  • Developers can access logs for debugging.

Export Summary file

As a developer or higher, you can generate an Export Summary to quickly understand the contents of your downloaded dataset, including item names, annotations, and metadata.

Learn more to export the summary file via UI and SDK.


Folders Structure

Datasets allows you to organize files in nested folders structure. Folder actions supported in the platform, via user-interface and SDK/API, are:

  • Create folder

  • Move item to folder (single or Bulk)

  • Clone item(s) to folder

  • Delete folder