BladePipe 1.9.0: New data pipelines, faster Oracle writes, and improved stability.
跳到主要内容

Oracle 到 Iceberg

选择对端数据库:

数据链路

基本功能

功能说明
Schema Migration

If the target schema does not exist, BladePipe will automatically generate and execute CREATE statements based on the source metadata and the mapping rule.

Full Data Migration

Migrate data by sequentially scanning data in tables and writing it in batches to the target database.

Incremental Data Sync

Sync of common DML like INSERT, UPDATE, DELETE is supported.
UPDATE and DELETE for tables without primary keys are not synced by default (manual selection required).

Subscription Modification

Add, delete, or modify the subscribed tables with support for historical data migration. For more information, see Modify Subscription.

Position Resetting

Reset the position by timestamp or Scn to consume Oracle Redo Log in a past period again.

Table Name Mapping

Support the mapping rules, namely, keeping the name the same as that in Source, converting the text to lowercase, converting the text to uppercase, truncating the name by "_digit" suffix.

DDL Synchronization

Supports ALTER TABLE ADD COLUMN and DROP COLUMN.

高级功能

功能说明
自动建字典

如果使用离线字典解析 Oracle Redo, 则在创建任务时自动创建字典

写入冲突策略

源端有主键表进行覆盖写入,源端无主键表进行追加写入

自定义表属性

包括 format-version 等属性设置

设置数据分区

创建任务时,可按表粒度指定分区定义(静态或动态),结构迁移时自动添加该分区定义

Custom Code

For more information, see Custom Code Processing, Debug Custom Code and Logging in Custom Code.

Setting Target Primary Key

Change the primary key to another field to facilitate data aggregation and other operations.

Data Filtering Conditions

Support data filtering using WHERE conditions, with SQL-92 as the SQL language. For more information, see Data Filtering.

限制和注意点

限制项说明
增量同步性能

因 Logminer 有性能上限,且 CloudCanal 未采用并行分析,所以以 3000 条变更/秒 为性能基准

数据类型

不支持 BLOB 及衍生类型

使用示例

标题详情
Oracle 数据迁移同步优化与思考

文档:Oracle 数据迁移同步优化与思考

Oracle 数据迁移同步优化(三)

文档:Oracle 数据迁移同步优化(三)


源端数据源

前置条件

条件说明
账号权限

文档:Oracle 需要的权限

增量同步准备

文档:Oracle Logminer 准备

网络准备

迁移同步节点(sidecar)可连接 ORACLE 标准交互接口(如 1521)

任务参数

参数名称说明
fullFetchSize

全量扫描数据设置的 fetch size

eventStoreSize

缓存解析完毕的增量事件缓存大小

logminerUser

执行 Logminer SQL 的 Oracle 连接用户

logminerPasswd

执行 Logminer SQL 的 Oracle 连接密码

logminerConnectType

执行 Logminer SQL 的 Oracle 连接类型(PDB),包括 ORACLE_SID, ORACLE_SERVICE 两种可选

logminerSidOrService

执行 Logminer SQL 的 Oracle 连接串 SID 或服务名(PDB)

parseRedoSqlParallel

解析 Logminer 数据的并发度

parseRedoSqlBufferSize

解析 Logminer 数据的环形队列大小

redoFetchSize

单次获取 Logminer 分析数据条数

redoOfferTransMaxSize

未消费但已提交事务最大缓存数量

oraMiningSessionPauseSec

使用 Logminer 挖掘日志间隙停顿时间,单位为秒

maxEventCountPerTxInMem

内存中每个事务的最大事件数

logMiningScnStep

Oracle Logminer 分析 redo log 时指定的分析范围大小

abandonUnCommitTxTimeoutSec

不带数据变更的事务未提交超过设置的值,则自动放弃该事务

restartTxWithDataTimeoutSec

带数据变更的事务未提交超过设置的值,则自动重启任务

oraUseOnlineDic

是否使用在线日志,false 使用离线日志对 Oracle 压力较大

oraReleaseIntervalSec

重建分析链接的间隔,以释放 Oracle 服务端资源

fallBackScnStep

和 Redo log 最新数据保持的距离,0 表示紧跟

sqlCaseConversionEnabled

是否打开 DDL 大小写转换(根据当前数据库默认大小写规则)

Tips: 通用参数配置请参考 通用参数及功能


目标端数据源

前置条件

条件说明
网络准备

迁移同步节点(sidecar)可连接 Catalog 和 文件存储

Nessie 数据源配置模版
  • 网络地址(CatalogUri): ip:19120/api/v1

  • catalogName: nessie

  • catalogType: NESSIE

  • catalogWarehouse: s3://warehouse

  • catalogProps: { "io-impl": "org.apache.iceberg.aws.s3.S3FileIO", "s3.endpoint": "http://ip:9000", "s3.access-key-id": "admin", "s3.secret-access-key": "password", "s3.path-style-access": "true", "client.region": "ap-southeast-1" }

Glue 数据源配置模版
  • 网络地址(CatalogUri): glue.ap-southeast-1.amazonaws.com

  • httpsEnabled: true

  • catalogName: glue_catalog

  • catalogType: GLUE

  • catalogWarehouse : s3://warehouse

  • catalogProps: { "io-impl": "org.apache.iceberg.aws.s3.S3FileIO", "s3.endpoint": "https://s3.ap-southeast-1.amazonaws.com", "s3.access-key-id": "", "s3.secret-access-key": "", "s3.path-style-access": "true", "client.region": "ap-southeast-1", "client.credentials-provider.glue.access-key-id": "", "client.credentials-provider.glue.secret-access-key": "", "client.credentials-provider": "com.amazonaws.glue.catalog.credentials.GlueAwsCredentialsProvider" }

REST 数据源配置模版
  • 网络地址(CatalogUri): ip :8181

  • httpsEnabled: false

  • catalogName: rest_catalog

  • catalogType: REST

  • catalogWarehouse : s3://warehouse

  • catalogProps: { "io-impl": "org.apache.iceberg.aws.s3.S3FileIO", "s3.endpoint": "http://ip:9000", "s3.access-key-id": "admin", "s3.secret-access-key": "password", "s3.path-style-access": "true", "client.region": "us-east-1" }

任务参数

参数名称说明
fileFormat

写入文件格式(parquet / orc / ... )

writeTargetFileSizeMb

写入目标文件大小(MB)

writeProps

写入配置参数(Json 格式)

commitBranch

写入提交的分支

totalDataInMemMb

攒批写入,内存中最大数据容量,超过此容量或超过 asyncFlushIntervalSec 则刷出数据到写入队列

asyncFlushIntervalSec

攒批写入,等待刷出的间隔时间,超过此时间或超过 totalDataInMemMb 则刷出数据到写入队列

flushBatchMb

单表最大攒批容量,超过此容量则刷出数据到写入队列

realFlushPauseSec

刷出数据到 Iceberg 的等待时间,0 则不等待

Tips: 通用参数配置请参考 通用参数及功能