DDOE delivers enterprise-grade unstructured data management and versioning capabilities, enabling sub-second queries across millions of files using item attributes, system metadata, and user-defined metadata.
The Data Management page provides a centralized workspace for managing datasets, storage drivers, preprocesses, and feature sets (embeddings), helping teams efficiently organize, analyze, and automate their data workflows.
.png)
Key Features
DDOE's Data Management provides a centralized platform for storing, organizing, processing, and analyzing data throughout the AI lifecycle. It enables teams to efficiently manage datasets, automate workflows, and prepare data for annotation, training, and model deployment.
Key capabilities include:
Flexible Dataset Creation: Upload data, sync cloud storage, connect on-premises storage, or use Compute Cluster storage integrations.
Dataset and File Management: Browse, filter, organize, move, clone, and delete files.
Task Integration: Create annotation tasks, launch pipelines, and trigger workflows from selected data.
Metadata Management: View metadata, logs, and import/export data.
Data Analysis: Identify duplicates, missing annotations, unlabeled data, and metadata issues.
Preprocess Automation: Run applications and pipelines automatically on new data.
Embeddings: Visualize and explore dataset embeddings.
Cloud and On-Premises Support: Connect AWS, Azure, GCP, and on-premises storage systems.
Advanced Querying (DQL): Search large datasets using metadata and attributes.
Developer Access: Access all features through APIs and SDKs.
Dedicated Views: Manage Datasets, Storage Drivers, Embeddings, and Preprocesses from dedicated tabs.
Datasets
DDOE enables you to create and manage datasets from a single location. You can work with internal storage, cloud storage providers, Compute Cluster storage integrations, and on-premises storage systems without switching between multiple interfaces. This centralized approach simplifies data access, organization, and management across your AI and data workflows.
Compute Clusters
A Compute Cluster connects DDOE to your Kubernetes environment and provides the infrastructure required to run applications, services, pipelines, preprocesses, and AI workloads. It serves as the execution layer where DDOE performs data processing, automation, and machine learning tasks.
Storage Integrations
Integrations securely connect DDOE to external data sources and storage providers such as AWS, Azure, Google Cloud Platform (GCP), and On-Premises systems, enabling flexible and seamless access to data.
Preprocesses
Preprocesses are automated services and pipelines that run when new items are created in project datasets. They help prepare, transform, enrich, or analyze data before it is used in downstream workflows such as annotation, training, and data processing.
Embeddings
The Embeddings in the DDOE platform provides a powerful way to visualize and interact with feature sets derived from your datasets. These feature sets represent the data in a numerical form (embeddings) that models can process and learn from.
Cloud Providers & Features
Cloud Provider | Resource Type | Integration Type |
|---|---|---|
AWS | S3 Bucket | Cross Account |
AWS | S3 Bucket | Access Key |
AWS | S3 Bucket | STS |
GCP | GCS Bucket | Private Key |
GCP | GCS Bucket | Cross Project |
Azure | Blob | Client Secret |
Azure | Datalake Gen2 | Client Secret |
DDOE supports sub-folder specific access in buckets, which offers security and versatility in managing your data. Learn more about the Specifications.