DOCS#
Tutorial
docs
- Operator Schemas 算子提要
- Dataset Configuration Guide
- “Bad” Data Exhibition
- Cache Management
- DJ-SORA
- DJ_service
- How-to Guide for Developers
- Distributed Data Processing in Data-Juicer
- Dataset Export
- Job Management
- Operator Plugins
- Partitioned Processing with Checkpointing
- Data Tracing
- Awesome Data-Model Co-Development of MLLMs
demos
- Demos
- Agent 交互数据:Bad case 洞察
- Bad case 自助报告(简化入口)
- Agent 流水线里 LLM 算子:加速与超参
- Bad case 流水线:一键运行与端到端指南
- Agent quality & bad-case docs
- Agent 质检 / bad-case 文档索引
- Agent 流水线最小可运行配置(便于逐项调试)
- Agent pipeline 后分析脚本
- VLA Visualization Demo
- Elastic Multi-Node Sharding on Shared Storage
- Overview
- Key advantages
- Comparison with existing approaches
- Files
- Important path concepts
- Requirements and current scope
- PAI-DLC Worker-broadcast quick start
- GPU smoke test: one recipe with CPU and GPU operators
- Important: use an existing Mapper/Filter recipe and your own JSONL
dlc_job.pyparameter reference- Complete
shard_job.pyparameter reference - Generic manual or scheduler workflow
- Ray execution modes
- Job directory layout
- Failure semantics
- Exit codes
- Troubleshooting
- Tests
- Note for dataset path
- HumanVBench Operators Demo
tools
- Distributed Fuzzy Deduplication Tools
- Auto Evaluation Toolkit
- GPT EVAL: Evaluate your model with OpenAI API
- Evaluation Results Recorder
- Format Conversion Tools
- Multimodal Tools
- Post Tuning Tools
- Label Studio Service Utility
- Metrics for video generation
- VBench metrics
- Postprocess tools
- Preprocess Tools