data-juicer data-juicer
Doc API
Sandbox Hub Agents
English 简体中文
main v1.5.5 v1.5.4 v1.5.3 v1.5.2 v1.5.1 v1.5.0 v1.4.6 v1.4.5 v1.4.4 v1.4.3 v1.4.2 v1.4.1 v1.4.0
  • Home

Documentation

Tutorial

  • DJ-Cookbook
  • Installation Guide
  • Quick Start

docs

  • Operator Schemas 算子提要
  • Dataset Configuration Guide
  • “Bad” Data Exhibition
  • Cache Management
  • DJ-SORA
  • DJ_service
  • How-to Guide for Developers
  • Distributed Data Processing in Data-Juicer
  • Dataset Export
  • Job Management
  • Partitioned Processing with Checkpointing
  • Data Tracing
  • Awesome Data-Model Co-Development of MLLMs

operators

  • Aggregator
    • entity_attribute_aggregator
    • meta_tags_aggregator
    • most_relevant_entities_aggregator
    • nested_aggregator
  • Deduplicator
    • document_deduplicator
    • document_minhash_deduplicator
    • document_minhash_deduplicator_with_uid
    • document_simhash_deduplicator
    • image_deduplicator
    • ray_bts_minhash_cpp_deduplicator
    • ray_bts_minhash_deduplicator
    • ray_bts_minhash_deduplicator_with_uid
    • ray_document_deduplicator
    • ray_image_deduplicator
    • ray_video_deduplicator
    • video_deduplicator
  • Filter
    • alphanumeric_filter
    • audio_duration_filter
    • audio_nmf_snr_filter
    • audio_size_filter
    • average_line_length_filter
    • character_repetition_filter
    • flagged_words_filter
    • general_field_filter
    • image_aesthetics_filter
    • image_aspect_ratio_filter
    • image_face_count_filter
    • image_face_ratio_filter
    • image_nsfw_filter
    • image_pair_similarity_filter
    • image_shape_filter
    • image_size_filter
    • image_subplot_filter
    • image_text_matching_filter
    • image_text_similarity_filter
    • image_watermark_filter
    • in_context_influence_filter
    • instruction_following_difficulty_filter
    • language_id_score_filter
    • llm_analysis_filter
    • llm_condition_filter
    • llm_difficulty_score_filter
    • llm_perplexity_filter
    • llm_quality_score_filter
    • llm_task_relevance_filter
    • maximum_line_length_filter
    • perplexity_filter
    • phrase_grounding_recall_filter
    • special_characters_filter
    • specified_field_filter
    • specified_numeric_field_filter
    • stopwords_filter
    • suffix_filter
    • text_action_filter
    • text_embd_similarity_filter
    • text_entity_dependency_filter
    • text_length_filter
    • text_pair_similarity_filter
    • token_num_filter
    • video_aesthetics_filter
    • video_aspect_ratio_filter
    • video_duration_filter
    • video_frames_text_similarity_filter
    • video_motion_score_filter
    • video_motion_score_ptlflow_filter
    • video_motion_score_raft_filter
    • video_nsfw_filter
    • video_ocr_area_ratio_filter
    • video_resolution_filter
    • video_tagging_from_frames_filter
    • video_watermark_filter
    • word_repetition_filter
    • words_num_filter
  • Mapper
    • audio_add_gaussian_noise_mapper
    • audio_ffmpeg_wrapped_mapper
    • calibrate_qa_mapper
    • calibrate_query_mapper
    • calibrate_response_mapper
    • chinese_convert_mapper
    • clean_copyright_mapper
    • clean_email_mapper
    • clean_html_mapper
    • clean_ip_mapper
    • clean_links_mapper
    • detect_character_attributes_mapper
    • detect_character_locations_mapper
    • detect_main_character_mapper
    • dialog_intent_detection_mapper
    • dialog_sentiment_detection_mapper
    • dialog_sentiment_intensity_mapper
    • dialog_topic_detection_mapper
    • download_file_mapper
    • expand_macro_mapper
    • extract_entity_attribute_mapper
    • extract_entity_relation_mapper
    • extract_event_mapper
    • extract_keyword_mapper
    • extract_nickname_mapper
    • extract_support_text_mapper
    • extract_tables_from_html_mapper
    • fix_unicode_mapper
    • general_fused_op
    • generate_qa_from_examples_mapper
    • generate_qa_from_text_mapper
    • human_preference_annotation_mapper
    • image_blur_mapper
    • image_captioning_mapper
    • image_detection_yolo_mapper
    • image_diffusion_mapper
    • image_face_blur_mapper
    • image_mmpose_mapper
    • image_remove_background_mapper
    • image_sam_3d_body_mapper
    • image_segment_mapper
    • image_tagging_mapper
    • image_tagging_vlm_mapper
    • imgdiff_difference_area_generator_mapper
    • imgdiff_difference_caption_generator_mapper
    • latex_figure_context_extractor_mapper
    • latex_merge_tex_mapper
    • llm_extract_mapper
    • mllm_mapper
    • nlpaug_en_mapper
    • nlpcda_zh_mapper
    • optimize_prompt_mapper
    • optimize_qa_mapper
    • optimize_query_mapper
    • optimize_response_mapper
    • pair_preference_mapper
    • punctuation_normalization_mapper
    • python_file_mapper
    • python_lambda_mapper
    • query_intent_detection_mapper
    • query_sentiment_detection_mapper
    • query_topic_detection_mapper
    • relation_identity_mapper
    • remove_bibliography_mapper
    • remove_comments_mapper
    • remove_header_mapper
    • remove_long_words_mapper
    • remove_non_chinese_character_mapper
    • remove_repeat_sentences_mapper
    • remove_specific_chars_mapper
    • remove_table_text_mapper
    • remove_words_with_incorrect_substrings_mapper
    • replace_content_mapper
    • s3_download_file_mapper
    • s3_upload_file_mapper
    • sdxl_prompt2prompt_mapper
    • sentence_augmentation_mapper
    • sentence_split_mapper
    • text_chunk_mapper
    • text_tagging_by_prompt_mapper
    • vggt_mapper
    • video_camera_calibration_static_deepcalib_mapper
    • video_camera_calibration_static_moge_mapper
    • video_captioning_from_audio_mapper
    • video_captioning_from_frames_mapper
    • video_captioning_from_summarizer_mapper
    • video_captioning_from_video_mapper
    • video_captioning_from_vlm_mapper
    • video_depth_estimation_mapper
    • video_extract_frames_mapper
    • video_face_blur_mapper
    • video_ffmpeg_wrapped_mapper
    • video_hand_reconstruction_mapper
    • video_object_segmenting_mapper
    • video_remove_watermark_mapper
    • video_resize_aspect_ratio_mapper
    • video_resize_resolution_mapper
    • video_split_by_duration_mapper
    • video_split_by_key_frame_mapper
    • video_split_by_scene_mapper
    • video_tagging_from_audio_mapper
    • video_tagging_from_frames_mapper
    • video_undistort_mapper
    • video_whole_body_pose_estimation_mapper
    • whitespace_normalization_mapper
  • Formatter
    • csv_formatter
    • empty_formatter
    • json_formatter
    • parquet_formatter
    • ray_empty_formatter
    • text_formatter
    • tsv_formatter
  • Grouper
    • key_value_grouper
    • naive_grouper
    • naive_reverse_grouper
  • Selector
    • frequency_specified_field_selector
    • random_selector
    • range_specified_field_selector
    • tags_specified_field_selector
    • topk_specified_field_selector
  • Pipeline
    • llm_ray_vllm_engine_pipeline
    • vlm_ray_vllm_engine_pipeline

demos

  • Demos
  • Agent 交互数据:Bad case 洞察
  • Bad case 自助报告(简化入口)
  • Agent 流水线里 LLM 算子:加速与超参
  • Bad case 流水线:一键运行与端到端指南
  • Agent quality & bad-case docs
  • Agent 质检 / bad-case 文档索引
  • Agent 流水线最小可运行配置(便于逐项调试)
  • Agent pipeline 后分析脚本
  • Note for dataset path

tools

  • Distributed Fuzzy Deduplication Tools
  • Auto Evaluation Toolkit
  • GPT EVAL: Evaluate your model with OpenAI API
  • Evaluation Results Recorder
  • Format Conversion Tools
  • Multimodal Tools
  • Post Tuning Tools
  • Label Studio Service Utility
  • Metrics for video generation
  • VBench metrics
  • Postprocess tools
  • Preprocess Tools

thirdparty

  • LLM Ecosystems
  • Third-party Model Library

Related projects

Sandbox Hub Agents

data_juicer.ops.mapper.clean_copyright_mapper module¶

class data_juicer.ops.mapper.clean_copyright_mapper.CleanCopyrightMapper(*args, **kwargs)[source]¶

Bases: Mapper

Cleans copyright comments at the beginning of text samples.

This operator removes copyright comments from the start of text samples. It identifies and strips multiline comments that contain the word “copyright” using a regular expression. It also greedily removes lines starting with comment markers like //, #, or – at the beginning of the text, as these are often part of copyright headers. The operator processes each sample individually but can handle batches for efficiency.

__init__(*args, **kwargs)[source]¶

Initialization method.

Parameters:
  • args – extra args

  • kwargs – extra args

process_batched(samples)[source]¶

On this page

  • data_juicer.ops.mapper.clean_copyright_mapper module
    • CleanCopyrightMapper
      • CleanCopyrightMapper.__init__()
      • CleanCopyrightMapper.process_batched()

© 2024, Data-Juicer Team — Built with Sphinx & Data-Juicer Theme

↵ to select ↑↓ to navigate esc to close