data-juicer data-juicer
Doc API
Sandbox Hub Agents
English 简体中文
main v1.5.5 v1.5.4 v1.5.3 v1.5.2 v1.5.1 v1.5.0 v1.4.6 v1.4.5 v1.4.4 v1.4.3 v1.4.2 v1.4.1 v1.4.0
  • 首页

文档

教程

  • DJ-Cookbook
  • 安装
  • 快速上手

帮助文档

  • Operator Schemas 算子提要
  • 数据集配置指南
  • “坏”数据展览
  • 缓存管理
  • DJ-SORA
  • DJ_服务化
  • 开发者指南
  • Data-Juicer 分布式数据处理
  • 数据集导出
  • 作业管理
  • 分区处理与检查点
  • 数据追踪
  • Awesome Data-Model Co-Development of MLLMs

算子

  • Aggregator
    • entity_attribute_aggregator
    • meta_tags_aggregator
    • most_relevant_entities_aggregator
    • nested_aggregator
  • Deduplicator
    • document_deduplicator
    • document_minhash_deduplicator
    • document_minhash_deduplicator_with_uid
    • document_simhash_deduplicator
    • image_deduplicator
    • ray_bts_minhash_cpp_deduplicator
    • ray_bts_minhash_deduplicator
    • ray_bts_minhash_deduplicator_with_uid
    • ray_document_deduplicator
    • ray_image_deduplicator
    • ray_video_deduplicator
    • video_deduplicator
  • Filter
    • alphanumeric_filter
    • audio_duration_filter
    • audio_nmf_snr_filter
    • audio_size_filter
    • average_line_length_filter
    • character_repetition_filter
    • flagged_words_filter
    • general_field_filter
    • image_aesthetics_filter
    • image_aspect_ratio_filter
    • image_face_count_filter
    • image_face_ratio_filter
    • image_nsfw_filter
    • image_pair_similarity_filter
    • image_shape_filter
    • image_size_filter
    • image_subplot_filter
    • image_text_matching_filter
    • image_text_similarity_filter
    • image_watermark_filter
    • in_context_influence_filter
    • instruction_following_difficulty_filter
    • language_id_score_filter
    • llm_analysis_filter
    • llm_condition_filter
    • llm_difficulty_score_filter
    • llm_perplexity_filter
    • llm_quality_score_filter
    • llm_task_relevance_filter
    • maximum_line_length_filter
    • perplexity_filter
    • phrase_grounding_recall_filter
    • special_characters_filter
    • specified_field_filter
    • specified_numeric_field_filter
    • stopwords_filter
    • suffix_filter
    • text_action_filter
    • text_embd_similarity_filter
    • text_entity_dependency_filter
    • text_length_filter
    • text_pair_similarity_filter
    • token_num_filter
    • video_aesthetics_filter
    • video_aspect_ratio_filter
    • video_duration_filter
    • video_face_ratio_filter
    • video_frames_text_similarity_filter
    • video_motion_score_filter
    • video_motion_score_ptlflow_filter
    • video_motion_score_raft_filter
    • video_nsfw_filter
    • video_ocr_area_ratio_filter
    • video_resolution_filter
    • video_tagging_from_frames_filter
    • video_watermark_filter
    • word_repetition_filter
    • words_num_filter
  • Mapper
    • audio_add_gaussian_noise_mapper
    • audio_ffmpeg_wrapped_mapper
    • calibrate_qa_mapper
    • calibrate_query_mapper
    • calibrate_response_mapper
    • chinese_convert_mapper
    • clean_copyright_mapper
    • clean_email_mapper
    • clean_html_mapper
    • clean_ip_mapper
    • clean_links_mapper
    • detect_character_attributes_mapper
    • detect_character_locations_mapper
    • detect_main_character_mapper
    • dialog_intent_detection_mapper
    • dialog_sentiment_detection_mapper
    • dialog_sentiment_intensity_mapper
    • dialog_topic_detection_mapper
    • download_file_mapper
    • expand_macro_mapper
    • extract_entity_attribute_mapper
    • extract_entity_relation_mapper
    • extract_event_mapper
    • extract_keyword_mapper
    • extract_nickname_mapper
    • extract_support_text_mapper
    • extract_tables_from_html_mapper
    • fix_unicode_mapper
    • general_fused_op
    • generate_qa_from_examples_mapper
    • generate_qa_from_text_mapper
    • human_preference_annotation_mapper
    • image_blur_mapper
    • image_captioning_mapper
    • image_detection_yolo_mapper
    • image_diffusion_mapper
    • image_face_blur_mapper
    • image_mmpose_mapper
    • image_remove_background_mapper
    • image_sam_3d_body_mapper
    • image_segment_mapper
    • image_tagging_mapper
    • image_tagging_vlm_mapper
    • imgdiff_difference_area_generator_mapper
    • imgdiff_difference_caption_generator_mapper
    • latex_figure_context_extractor_mapper
    • latex_merge_tex_mapper
    • llm_extract_mapper
    • mllm_mapper
    • nlpaug_en_mapper
    • nlpcda_zh_mapper
    • optimize_prompt_mapper
    • optimize_qa_mapper
    • optimize_query_mapper
    • optimize_response_mapper
    • pair_preference_mapper
    • punctuation_normalization_mapper
    • python_file_mapper
    • python_lambda_mapper
    • query_intent_detection_mapper
    • query_sentiment_detection_mapper
    • query_topic_detection_mapper
    • relation_identity_mapper
    • remove_bibliography_mapper
    • remove_comments_mapper
    • remove_header_mapper
    • remove_long_words_mapper
    • remove_non_chinese_character_mapper
    • remove_repeat_sentences_mapper
    • remove_specific_chars_mapper
    • remove_table_text_mapper
    • remove_words_with_incorrect_substrings_mapper
    • replace_content_mapper
    • s3_download_file_mapper
    • s3_upload_file_mapper
    • sdxl_prompt2prompt_mapper
    • sentence_augmentation_mapper
    • sentence_split_mapper
    • text_chunk_mapper
    • text_tagging_by_prompt_mapper
    • vggt_mapper
    • video_active_speaker_detect_mapper
    • video_audio_ASR_mapper
    • video_audio_detect_age_gender_mapper
    • video_audio_speech_emotion_mapper
    • video_camera_calibration_deepcalib_mapper
    • video_camera_calibration_moge_mapper
    • video_captioning_face_attribute_emotion_mapper
    • video_captioning_from_audio_mapper
    • video_captioning_from_frames_mapper
    • video_captioning_from_human_tracks_mapper
    • video_captioning_from_summarizer_mapper
    • video_captioning_from_video_mapper
    • video_captioning_from_vlm_mapper
    • video_depth_estimation_mapper
    • video_extract_frames_mapper
    • video_face_blur_mapper
    • video_ffmpeg_wrapped_mapper
    • video_hand_reconstruction_mapper
    • video_human_tracks_extraction_mapper
    • video_human_tracks_face_demographic_mapper
    • video_object_segmenting_mapper
    • video_remove_watermark_mapper
    • video_resize_aspect_ratio_mapper
    • video_resize_resolution_mapper
    • video_split_by_duration_mapper
    • video_split_by_key_frame_mapper
    • video_split_by_scene_mapper
    • video_tagging_from_audio_mapper
    • video_tagging_from_frames_mapper
    • video_undistort_mapper
    • video_whole_body_pose_estimation_mapper
    • whitespace_normalization_mapper
  • Formatter
    • csv_formatter
    • empty_formatter
    • json_formatter
    • parquet_formatter
    • ray_empty_formatter
    • text_formatter
    • tsv_formatter
  • Grouper
    • key_value_grouper
    • naive_grouper
    • naive_reverse_grouper
  • Selector
    • frequency_specified_field_selector
    • random_selector
    • range_specified_field_selector
    • tags_specified_field_selector
    • topk_specified_field_selector
  • Pipeline
    • llm_ray_vllm_engine_pipeline
    • ray_repartition_pipeline
    • vlm_ray_vllm_engine_pipeline

demos

  • 演示
  • Agent 交互数据:Bad case 洞察
  • Bad case 自助报告(简化入口)
  • Agent 流水线里 LLM 算子:加速与超参
  • Bad case 流水线:一键运行与端到端指南
  • Agent quality & bad-case docs
  • Agent 质检 / bad-case 文档索引
  • Agent 流水线最小可运行配置(便于逐项调试)
  • Agent pipeline 后分析脚本
  • 自动化评测:HELM 评测及可视化
  • VLA Visualization Demo
  • Note for dataset path
  • 为LLM构造角色扮演的system prompt
  • HumanVBench 算子演示

工具

  • 分布式模糊去重工具
  • Auto Evaluation Toolkit
  • GPT EVAL:使用 OpenAI API 评测大模型
  • Evaluation Results Recorder
  • 格式转换工具
  • 多模态工具
  • 后微调工具
  • Label Studio Service Utility
  • 视频生成测评工具
  • VBench metrics
  • Postprocess tools
  • 预处理工具

第三方

  • 大语言模型生态
  • HumanVBench Models Setup
  • 第三方模型库

相关项目

Sandbox Hub Agents

data_juicer.ops.mapper.clean_copyright_mapper module¶

class data_juicer.ops.mapper.clean_copyright_mapper.CleanCopyrightMapper(*args, **kwargs)[源代码]¶

基类:Mapper

Cleans copyright comments at the beginning of text samples.

This operator removes copyright comments from the start of text samples. It identifies and strips multiline comments that contain the word "copyright" using a regular expression. It also greedily removes lines starting with comment markers like //, #, or -- at the beginning of the text, as these are often part of copyright headers. The operator processes each sample individually but can handle batches for efficiency.

__init__(*args, **kwargs)[源代码]¶

Initialization method.

参数:
  • args -- extra args

  • kwargs -- extra args

process_batched(samples)[源代码]¶

On this page

  • data_juicer.ops.mapper.clean_copyright_mapper module
    • CleanCopyrightMapper
      • CleanCopyrightMapper.__init__()
      • CleanCopyrightMapper.process_batched()

© 2024, Data-Juicer Team — Built with Sphinx & Data-Juicer Theme

↵ to select ↑↓ to navigate esc to close