data-juicer data-juicer
Doc API
Sandbox Hub Agents
English 简体中文
main v1.5.5 v1.5.4 v1.5.3 v1.5.2 v1.5.1 v1.5.0 v1.4.6 v1.4.5 v1.4.4 v1.4.3 v1.4.2 v1.4.1 v1.4.0
  • Home

Documentation

Tutorial

  • DJ-Cookbook
  • Installation Guide
  • Quick Start

docs

  • Operator Schemas 算子提要
  • Dataset Configuration Guide
  • “Bad” Data Exhibition
  • Cache Management
  • DJ-SORA
  • DJ_service
  • How-to Guide for Developers
  • Distributed Data Processing in Data-Juicer
  • Dataset Export
  • Job Management
  • Partitioned Processing with Checkpointing
  • Data Tracing
  • Awesome Data-Model Co-Development of MLLMs

operators

  • Aggregator
    • entity_attribute_aggregator
    • meta_tags_aggregator
    • most_relevant_entities_aggregator
    • nested_aggregator
  • Deduplicator
    • document_deduplicator
    • document_minhash_deduplicator
    • document_minhash_deduplicator_with_uid
    • document_simhash_deduplicator
    • image_deduplicator
    • ray_bts_minhash_cpp_deduplicator
    • ray_bts_minhash_deduplicator
    • ray_bts_minhash_deduplicator_with_uid
    • ray_document_deduplicator
    • ray_image_deduplicator
    • ray_video_deduplicator
    • video_deduplicator
  • Filter
    • alphanumeric_filter
    • audio_duration_filter
    • audio_nmf_snr_filter
    • audio_size_filter
    • average_line_length_filter
    • character_repetition_filter
    • flagged_words_filter
    • general_field_filter
    • image_aesthetics_filter
    • image_aspect_ratio_filter
    • image_face_count_filter
    • image_face_ratio_filter
    • image_nsfw_filter
    • image_pair_similarity_filter
    • image_shape_filter
    • image_size_filter
    • image_subplot_filter
    • image_text_matching_filter
    • image_text_similarity_filter
    • image_watermark_filter
    • in_context_influence_filter
    • instruction_following_difficulty_filter
    • language_id_score_filter
    • llm_analysis_filter
    • llm_condition_filter
    • llm_difficulty_score_filter
    • llm_perplexity_filter
    • llm_quality_score_filter
    • llm_task_relevance_filter
    • maximum_line_length_filter
    • perplexity_filter
    • phrase_grounding_recall_filter
    • special_characters_filter
    • specified_field_filter
    • specified_numeric_field_filter
    • stopwords_filter
    • suffix_filter
    • text_action_filter
    • text_embd_similarity_filter
    • text_entity_dependency_filter
    • text_length_filter
    • text_pair_similarity_filter
    • token_num_filter
    • video_aesthetics_filter
    • video_aspect_ratio_filter
    • video_duration_filter
    • video_frames_text_similarity_filter
    • video_motion_score_filter
    • video_motion_score_ptlflow_filter
    • video_motion_score_raft_filter
    • video_nsfw_filter
    • video_ocr_area_ratio_filter
    • video_resolution_filter
    • video_tagging_from_frames_filter
    • video_watermark_filter
    • word_repetition_filter
    • words_num_filter
  • Mapper
    • audio_add_gaussian_noise_mapper
    • audio_ffmpeg_wrapped_mapper
    • calibrate_qa_mapper
    • calibrate_query_mapper
    • calibrate_response_mapper
    • chinese_convert_mapper
    • clean_copyright_mapper
    • clean_email_mapper
    • clean_html_mapper
    • clean_ip_mapper
    • clean_links_mapper
    • detect_character_attributes_mapper
    • detect_character_locations_mapper
    • detect_main_character_mapper
    • dialog_intent_detection_mapper
    • dialog_sentiment_detection_mapper
    • dialog_sentiment_intensity_mapper
    • dialog_topic_detection_mapper
    • download_file_mapper
    • expand_macro_mapper
    • extract_entity_attribute_mapper
    • extract_entity_relation_mapper
    • extract_event_mapper
    • extract_keyword_mapper
    • extract_nickname_mapper
    • extract_support_text_mapper
    • extract_tables_from_html_mapper
    • fix_unicode_mapper
    • general_fused_op
    • generate_qa_from_examples_mapper
    • generate_qa_from_text_mapper
    • human_preference_annotation_mapper
    • image_blur_mapper
    • image_captioning_mapper
    • image_detection_yolo_mapper
    • image_diffusion_mapper
    • image_face_blur_mapper
    • image_mmpose_mapper
    • image_remove_background_mapper
    • image_sam_3d_body_mapper
    • image_segment_mapper
    • image_tagging_mapper
    • image_tagging_vlm_mapper
    • imgdiff_difference_area_generator_mapper
    • imgdiff_difference_caption_generator_mapper
    • latex_figure_context_extractor_mapper
    • latex_merge_tex_mapper
    • llm_extract_mapper
    • mllm_mapper
    • nlpaug_en_mapper
    • nlpcda_zh_mapper
    • optimize_prompt_mapper
    • optimize_qa_mapper
    • optimize_query_mapper
    • optimize_response_mapper
    • pair_preference_mapper
    • punctuation_normalization_mapper
    • python_file_mapper
    • python_lambda_mapper
    • query_intent_detection_mapper
    • query_sentiment_detection_mapper
    • query_topic_detection_mapper
    • relation_identity_mapper
    • remove_bibliography_mapper
    • remove_comments_mapper
    • remove_header_mapper
    • remove_long_words_mapper
    • remove_non_chinese_character_mapper
    • remove_repeat_sentences_mapper
    • remove_specific_chars_mapper
    • remove_table_text_mapper
    • remove_words_with_incorrect_substrings_mapper
    • replace_content_mapper
    • s3_download_file_mapper
    • s3_upload_file_mapper
    • sdxl_prompt2prompt_mapper
    • sentence_augmentation_mapper
    • sentence_split_mapper
    • text_chunk_mapper
    • text_tagging_by_prompt_mapper
    • vggt_mapper
    • video_camera_calibration_static_deepcalib_mapper
    • video_camera_calibration_static_moge_mapper
    • video_captioning_from_audio_mapper
    • video_captioning_from_frames_mapper
    • video_captioning_from_summarizer_mapper
    • video_captioning_from_video_mapper
    • video_captioning_from_vlm_mapper
    • video_depth_estimation_mapper
    • video_extract_frames_mapper
    • video_face_blur_mapper
    • video_ffmpeg_wrapped_mapper
    • video_hand_reconstruction_mapper
    • video_object_segmenting_mapper
    • video_remove_watermark_mapper
    • video_resize_aspect_ratio_mapper
    • video_resize_resolution_mapper
    • video_split_by_duration_mapper
    • video_split_by_key_frame_mapper
    • video_split_by_scene_mapper
    • video_tagging_from_audio_mapper
    • video_tagging_from_frames_mapper
    • video_undistort_mapper
    • video_whole_body_pose_estimation_mapper
    • whitespace_normalization_mapper
  • Formatter
    • csv_formatter
    • empty_formatter
    • json_formatter
    • parquet_formatter
    • ray_empty_formatter
    • text_formatter
    • tsv_formatter
  • Grouper
    • key_value_grouper
    • naive_grouper
    • naive_reverse_grouper
  • Selector
    • frequency_specified_field_selector
    • random_selector
    • range_specified_field_selector
    • tags_specified_field_selector
    • topk_specified_field_selector
  • Pipeline
    • llm_ray_vllm_engine_pipeline
    • vlm_ray_vllm_engine_pipeline

demos

  • Demos
  • Agent 交互数据:Bad case 洞察
  • Bad case 自助报告(简化入口)
  • Agent 流水线里 LLM 算子:加速与超参
  • Bad case 流水线:一键运行与端到端指南
  • Agent quality & bad-case docs
  • Agent 质检 / bad-case 文档索引
  • Agent 流水线最小可运行配置(便于逐项调试)
  • Agent pipeline 后分析脚本
  • Note for dataset path

tools

  • Distributed Fuzzy Deduplication Tools
  • Auto Evaluation Toolkit
  • GPT EVAL: Evaluate your model with OpenAI API
  • Evaluation Results Recorder
  • Format Conversion Tools
  • Multimodal Tools
  • Post Tuning Tools
  • Label Studio Service Utility
  • Metrics for video generation
  • VBench metrics
  • Postprocess tools
  • Preprocess Tools

thirdparty

  • LLM Ecosystems
  • Third-party Model Library

Related projects

Sandbox Hub Agents

clean_links_mapper¶

Mapper to clean links like http/https/ftp in text samples.

This operator removes or replaces URLs and other web links in the text. It uses a regular expression pattern to identify and remove links. By default, it replaces the identified links with an empty string, effectively removing them. The operator can be customized with a different pattern and replacement string. It processes samples in batches and modifies the text in place. If no links are found in a sample, it is left unchanged.

映射器用于清理文本样本中的http/https/ftp等链接。

此算子删除或替换文本中的URL和其他网络链接。它使用正则表达式模式来识别和删除链接。默认情况下,它将识别到的链接替换为空字符串,从而删除它们。可以通过不同的模式和替换字符串自定义算子。它以批量方式处理样本并在原地修改文本。如果样本中没有找到链接,则保持不变。

Type 算子类型: mapper

Tags 标签: cpu, text

🔧 Parameter Configuration 参数配置¶

name 参数名

type 类型

default 默认值

desc 说明

pattern

typing.Optional[str]

None

regular expression pattern to search for within text.

repl

<class ‘str’>

''

replacement string, default is empty string.

args

''

extra args

kwargs

''

extra args

📊 Effect demonstration 效果演示¶

test_mixed_https_links_text¶

CleanLinksMapper()

📥 input data 输入数据¶

Sample 1: text
This is a test,https://www.example.com/file.html?param1=value1&param2=value2
Sample 2: text
这是个测试,https://example.com/my-page.html?param1=value1&param2=value2
Sample 3: text
这是个测试,https://example.com

📤 output data 输出数据¶

Sample 1: text
This is a test,
Sample 2: text
这是个测试,
Sample 3: text
这是个测试,

✨ explanation 解释¶

This example shows the operator removing HTTPS links from text that contains both plain text and a link. The operator identifies and removes the links, leaving the rest of the text intact. For example, ‘This is a test,https://www.example.com/file.html?param1=value1&param2=value2’ becomes ‘This is a test,’ after processing. 这个示例展示了算子从同时包含纯文本和链接的文本中移除HTTPS链接。算子识别并移除这些链接,而保留其余文本不变。例如,’This is a test,https://www.example.com/file.html?param1=value1&param2=value2’ 在处理后变为 ‘This is a test,’。

test_replace_links_text¶

CleanLinksMapper(repl='<LINKS>')

📥 input data 输入数据¶

Sample 1: text
ftp://user:password@ftp.example.com:21/
Sample 2: text
This is a sample for test
Sample 3: text
abcd://ef is a sample for test
Sample 4: text
HTTP://example.com/my-page.html?param1=value1&param2=value2

📤 output data 输出数据¶

Sample 1: text
<LINKS>
Sample 2: text
This is a sample for test
Sample 3: text
<LINKS> is a sample for test
Sample 4: text
<LINKS>

✨ explanation 解释¶

This example demonstrates the operator replacing different types of links with a custom string ‘’. If a sample contains a link, it will be replaced by ‘’, while samples without links remain unchanged. For instance, ‘ftp://user:password@ftp.example.com:21/’ is transformed into ‘’, whereas ‘This is a sample for test’ stays as it is because it doesn’t contain any links. 这个示例展示了算子使用自定义字符串’’替换不同类型的链接。如果一个样本包含链接,它将被替换为’’,而不含链接的样本则保持不变。例如,’ftp://user:password@ftp.example.com:21/’ 被转换为 ‘’,而 ‘This is a sample for test’ 保持不变,因为它不包含任何链接。

🔗 related links 相关链接¶

  • source code 源代码

  • unit test 单元测试

  • Return operator list 返回算子列表

On this page

  • clean_links_mapper
    • 🔧 Parameter Configuration 参数配置
    • 📊 Effect demonstration 效果演示
      • test_mixed_https_links_text
        • 📥 input data 输入数据
        • 📤 output data 输出数据
        • ✨ explanation 解释
      • test_replace_links_text
        • 📥 input data 输入数据
        • 📤 output data 输出数据
        • ✨ explanation 解释
    • 🔗 related links 相关链接
Previous clean_ip_mapper Next detect_character_attributes_mapper

© 2024, Data-Juicer Team — Built with Sphinx & Data-Juicer Theme

↵ to select ↑↓ to navigate esc to close