BladePipe 1.9.0: New data pipelines, faster Oracle writes, and improved stability.
跳到主要内容

SAP HANA 到 StarRocks

选择对端数据库:

数据链路

基本功能

功能说明
Schema Migration

If the target schema does not exist, BladePipe will automatically generate and execute CREATE statements based on the source metadata and the mapping rule.

Full Data Migration

Migrate data by sequentially scanning data in tables and writing it in batches to the target database.

Incremental Data Sync

Sync of common DML like INSERT, UPDATE, DELETE is supported.
UPDATE and DELETE for tables without primary keys are not synced by default (manual selection required).

Data Verification and Correction

Verify all existing data. Optionally, you can correct the inconsistent data based on verification results. Scheduled DataTasks are supported.
For more information, see Create Verification and Correction DataJob.

Subscription Modification

Add, delete, or modify the subscribed tables with support for historical data migration. For more information, see Modify Subscription.

Position Resetting

Reset positions by data ID or timestamp. Allow re-consumption of CDC data in a past period.

Table Name Mapping

Support the mapping rules, namely, keeping the name the same as that in Source, converting the text to lowercase, converting the text to uppercase, truncating the name by "_digit" suffix.

Metadata Retrieval

Retrieve the target metadata with filtering conditions or target primary keys set from the source table.

高级功能

功能说明
基于 Trigger 增量同步

任务会自动创建表的触发器,触发器能捕获数据的 INSERT / UPDATE / DELETE 事件并写入增量 CDC 数据表

全量前清空目标数据

运行全量任务前清除老数据,包括重跑任务、定时全量迁移都会触发此能力

重建目标表

运行全量任务前重建目标表,包括重跑任务、定时全量迁移都会触发此能力

Stream Load 数据写入

采用 Stream Load 到 StarRocks Be 写入数据, 默认攒批写入,可动态调节刷出数据节奏和批次大小

0 值时间处理

支持将 0 值时间设置成不同类型的值,防止写入对端报错

Custom Table Properties

Include settings for properties such as bucket count and replica count.

Setting Data Partitions

When creating a DataJob, specify partition definitions at the table level (static or dynamic). Automatically add these partition definitions during schema migration.

Scheduled Full Data Migration

For more information, see Create Scheduled Full Data DataJob.

Custom Code

For more information, see Custom Code Processing, Debug Custom Code and Logging in Custom Code.

Data Filtering Conditions

Support data filtering using WHERE conditions, with SQL-92 as the SQL language. For more information, see Data Filtering.

限制和注意点

限制项说明
DDL 变化处理方案

SAP HANA 源端通过触发器捕获增量数据,不支持 DDL 同步。若发生 DDL 变更,可参考文档:SAP HANA 源端表结构变更

HANA 增量同步数据类型

HANA 增量阶段,触发器不支持捕获 TEXTBIN_TEXTST_POINTST_GEOMETRY 类型的数据变更

对端表类型

仅支持 主键模型(Primary Key)

源端表类型

不支持 无主键表 迁移同步

DDL 同步报错
  • 同一张表连续几个 DDL 将报错(因 StarRocks 对端是异步 DDL)
  • 修改字段约束或者部分类型的 DDL 报错
  • 如遇到 DDL 报错,可在对端变更好表结构,然后通过设置任务参数跳过,文档:跳过 DDL 异常
增量写入冲突策略

Stream Load 写入以主键进行整行替换


源端数据源

前置条件

条件说明
账号权限

文档:HANA 需要的权限

任务参数

参数名称说明
sysTriggerDataSchema

触发器写入增量表 SCHEMA 名称

sysTriggerDataTable

触发器写入增量表 TABLE 名称

incrPagingCount

触发器增量同步每次查询数据总量

incrIdleSleepSecond

触发器的增量同步空闲时查询间隔(单位:秒)

incrScanIntervalMs

设置基于触发器的增量同步数据查询间隔(单位:毫秒)

autoCheckTriggerAndReInstall

任务启动时检查触发器状态并重新安装

triggerDataCleanEnabled

是否开启定时清理触发器增量表数据

triggerDataCleanIntervalMin

设置触发器增量表的清理间隔(单位:分钟)

triggerDataRetentionMin

设置触发器增量表数据的保留时间(单位:分钟)

dbHeartbeatEnable

配置对源端数据库是否开启心跳

needTriggerDataJsonEscape

是否对触发器增量表数据加转义符(\)

triggerDataJsonQuotation

自定义触发器增量表 JSON 数据引号

triggerParamBathSize

设置触发器模板中每个变量包含列的个数

fullBeforeImageEnabled

触发器是否记录所有列变更前的完整数据

Tips: 通用参数配置请参考 通用参数及功能


目标端数据源

前置条件

条件说明
账号权限

具备 SELECT, DDL 权限(可选)

网络准备

迁移同步节点(sidecar)可连接 StarRocks FE QueryPortFE/BE HttpPort

任务参数

参数名称说明
host

MySQL 协议交互链接,对应 StarRocks FE QueryPort

httpHost

StarRocks stream load 链接,对应 StarRocks FE/BE HttpPort

totalDataInMemMb

攒批写入,内存中最大数据容量,超过此容量或超过 asyncFlushIntervalSec 则刷出数据到写入队列

asyncFlushIntervalSec

攒批写入,等待刷出的间隔时间,超过此时间或超过 totalDataInMemMb 则刷出数据到写入队列

flushBatchMb

单表最大攒批容量,超过此容量则刷出数据到写入队列

realFlushPauseSec

使用 stream load 刷出数据到 StarRocks 的等待时间,0 则不等待

soTimeoutSec

在 QueryPort 执行操作时 tcp 超时链接 (so_timeout)

httpSoTimeoutSec

在 HttpPort 执行操作时 tcp 超时链接 (so_timeout)

enableTimeZoneProcess

是否对时间字段进行时区转换

timezone

目标端 StarRocks 时区,例如 +08:00 Asia/Shanghai America/New_York

maxInSizePerQuery

校验任务中,对端单次查询的最大 IN 条件值数量,大于该值会自动拆分多次查询

Tips: 通用参数配置请参考 通用参数及功能