Documentation Index

Fetch the complete documentation index at: https://docs.dataloop.ai/llms.txt

Use this file to discover all available pages before exploring further.

Overview

Prev Next

DDOE delivers enterprise-grade unstructured data management and versioning capabilities, enabling sub-second queries across millions of files using item attributes, system metadata, and user-defined metadata.

The Data Management page provides a centralized workspace for managing datasets, storage drivers, preprocesses, and feature sets (embeddings), helping teams efficiently organize, analyze, and automate their data workflows.


Key Features

DDOE's Data Management provides a centralized platform for storing, organizing, processing, and analyzing data throughout the AI lifecycle. It enables teams to efficiently manage datasets, automate workflows, and prepare data for annotation, training, and model deployment.

Key capabilities include:

  • Flexible Dataset Creation: Upload data, sync cloud storage, connect on-premises storage, or use Compute Cluster storage integrations.

  • Dataset and File Management: Browse, filter, organize, move, clone, and delete files.

  • Task Integration: Create annotation tasks, launch pipelines, and trigger workflows from selected data.

  • Metadata Management: View metadata, logs, and import/export data.

  • Data Analysis: Identify duplicates, missing annotations, unlabeled data, and metadata issues.

  • Preprocess Automation: Run applications and pipelines automatically on new data.

  • Embeddings: Visualize and explore dataset embeddings.

  • Cloud and On-Premises Support: Connect AWS, Azure, GCP, and on-premises storage systems.

  • Advanced Querying (DQL): Search large datasets using metadata and attributes.

  • Developer Access: Access all features through APIs and SDKs.

  • Dedicated Views: Manage Datasets, Storage Drivers, Embeddings, and Preprocesses from dedicated tabs.


Datasets

DDOE enables you to create and manage datasets from a single location. You can work with internal storage, cloud storage providers, Compute Cluster storage integrations, and on-premises storage systems without switching between multiple interfaces. This centralized approach simplifies data access, organization, and management across your AI and data workflows.

Learn more

Compute Clusters

A Compute Cluster connects DDOE to your Kubernetes environment and provides the infrastructure required to run applications, services, pipelines, preprocesses, and AI workloads. It serves as the execution layer where DDOE performs data processing, automation, and machine learning tasks.

Learn more

Storage Integrations

Integrations securely connect DDOE to external data sources and storage providers such as AWS, Azure, Google Cloud Platform (GCP), and On-Premises systems, enabling flexible and seamless access to data.

Learn more

Preprocesses

Preprocesses are automated services and pipelines that run when new items are created in project datasets. They help prepare, transform, enrich, or analyze data before it is used in downstream workflows such as annotation, training, and data processing.

Learn more

Embeddings

The Embeddings in the DDOE platform provides a powerful way to visualize and interact with feature sets derived from your datasets. These feature sets represent the data in a numerical form (embeddings) that models can process and learn from.

Learn more


Cloud Providers & Features

Cloud Provider

Resource Type

Integration Type

AWS

S3 Bucket

Cross Account

AWS

S3 Bucket

Access Key

AWS

S3 Bucket

STS

GCP

GCS Bucket

Private Key

GCP

GCS Bucket

Cross Project

Azure

Blob

Client Secret

Azure

Datalake Gen2

Client Secret

DDOE supports sub-folder specific access in buckets, which offers security and versatility in managing your data. Learn more about the Specifications.