BladePipe 1.9.0: New data pipelines, faster Oracle writes, and improved stability.
跳到主要内容

TiDB 到 Elasticsearch

选择对端数据库:

数据链路

基本功能

功能说明
Schema Migration

If the target index does not exist, BladePipe will create the index mapping in the target based on the source metadata and mapping rules.

Full Data Migration

Migrate data by sequentially scanning data in tables and writing it in batches to the target database.

Incremental Data Sync

Sync of common DML like INSERT, UPDATE, DELETE is supported.
UPDATE and DELETE for tables without primary keys are not synced.

Data Verification and Correction

Verify all existing data. Optionally, you can correct the inconsistent data based on verification results. Scheduled DataTasks are supported.
For more information, see Create Verification and Correction DataJob.

Subscription Modification

Add, delete, or modify the subscribed tables with support for historical data migration. For more information, see Modify Subscription.

Position Resetting

Reset positions by timestamp to consume again the incremental data that has not been collected as garbage by TiKV in a past period.

Index Name Mapping

Support the mapping rules, namely, concatenation with underscores (dataJobName_DB_SCHEMA_table), keeping the name the same as that in Source, converting the text to lowercase, converting the text to uppercase, truncating the name by "_digit" suffix.

DDL Sync
  • ALTER TABLE ADD COLUMN
Metadata Retrieval

Retrieve the target metadata with filtering conditions or target primary keys set from the source table.

高级功能

功能说明
全量前清空目标数据

运行全量任务前清除老数据,包括重跑任务、定时全量迁移都会触发此能力

重建目标表

运行全量任务前重建目标表,包括重跑任务、定时全量迁移都会触发此能力

ES 时间写入格式

以该字段的第一个时间格式写入 Elasticsearch,如果未设置时间格式,则使用 yyyy-MM-dd'T'HH:mm:ss 格式

ES 时区设置

只有当时间格式的时区为 ZZZZZ 时,才会将页面设置的时区写入到 Elasticsearch

Optional Fields in Indexing

By default, all fields are indexed. Specific fields can be excluded from indexing.

Field-level Analyzers

Allow selecting analyzers for string fields that are indexed. Support STANDARD (default), SIMPLE, and other common analyzers, with the option to specify custom analyzers.

Setting Index _id Field

By default, the _id field is a concatenation of the source primary key values. It can be changed to other field values.

Scheduled Full Data Migration

For more information, see Create Scheduled Full Data DataJob.

Custom Code

For more information, see Custom Code Processing, Debug Custom Code and Logging in Custom Code.

Data Filtering Conditions

Support data filtering using WHERE conditions, with SQL-92 as the SQL language. For more information, see Data Filtering.

使用示例

标题详情
Elasticsearch 对端同步技术详解

文档:Elasticsearch 对端同步技术详解


源端数据源

前置条件

条件说明
账号权限

文档:TiDB 需要的权限

PD节点网络连通

请确保 CloudCanal 各节点能正常与 PD 各节点通讯

  • telnet [PD节点IP] [PD节点端口号]
TiKV GC 回收频率

在 TiDB Server 中修改 GC 周期时间为 24小时 以上

  • set global tidb_gc_life_time = "24h0m0s";
TiKV 历史变更数据缓存

建议根据任务所需适当调整大小

  • old-value-cache-memory-quota:增量旧数据占用 TiKV 节点的内存的上限
  • sink-memory-quota:增量数据占用 TiKV 节点的内存的上限

任务参数

参数名称说明
printDetailLog

打印接收到的增量,常用于判断源端是否有增量数据推送

pdHost

任务请求的 PD 节点地址,格式为: [PD_IP]:[PD_PORT], 多个 PD 节点用 , 隔开
例: 127.0.0.1:2379,127.0.0.1:2380

cdcGrpcTimeout

任务与 PD 节点 gRpc 连接通道的超时时间,单位ms

cdcStubTimeout

gRpc 通道中的每个 stub 的超时时间,超过该时间会自动重新订阅,单位ms

fastFailKeywords

字符串数组,以逗号分隔,当异常信息中包含这些关键字时,任务不再尝试重连,直接重启。例如 DEADLINE_EXCEEDED 表示当 gRPC 超时异常时不再重连,直接重启任务

Tips: 通用参数配置请参考 通用参数及功能


目标端数据源

前置条件

条件说明
账号权限

具备索引的 create, delete, create_index, delete_index, read, write 权限

网络准备

迁移同步节点(sidecar)可连接 Elasticsearch 节点

任务参数

参数名称说明
maxBulkSizeMb

单表最大攒批容量,超过此容量则刷出数据到写入队列

totalDataInMemMb

攒批写入,内存中最大数据容量,超过此容量或超过 asyncFlushIntervalSec 则刷出数据到写入队列

asyncFlushIntervalSec

攒批写入,等待刷出的间隔时间,超过此时间或超过 totalDataInMemMb 则刷出数据到写入队列

realFlushPauseSec

使用 Bulk Write 刷出数据到 ElasticSearch 的等待时间,0 则不等待

pkSeparator

拼接 _id 的分隔符(字段数 > 1)

enableBulkSizeThreshold

启用批量写入模式(默认开启)

Tips: 通用参数配置请参考 通用参数及功能