【问题标题】:bdr_init_copy hangs indefinitelybdr_init_copy 无限期挂起
【发布时间】:2016-10-12 22:46:28
【问题描述】:

对 Postgresql 来说相当新,但必须设置复制。我选择了 BDR,它在本地演示中运行良好,但在分布式机器上开始出现问题,主要是因为我不知道我到底在做什么,我哭着睡着,渴望 MySQL。我已经让 BDR 在多台服务器上工作,几乎。当我跑步时:

SELECT bdr.bdr_node_join_wait_for_ready();

在它挂起的加入节点上。这在 DB2 和 DB3 上都会发生。 DB1 返回一个有效响应。研究这个我遇到了 bdr_init_copy 命令,它显然做了我一直在手工做的所有事情,然后是一些。所以尝试了一下。现在,当我跑步时:

/usr/lib/postgresql/9.4/bin/bdr_init_copy -d "host=192.168.1.10 dbname=demo3" --local-dbname="host=192.168.1.23 dbname=demo3" -n db2 -D bdr_data

我明白了

bdr_init_copy: starting ...
Getting remote server identification ...
Detected 2 BDR database(s) on remote server
Updating BDR configuration on the remote node:
 demo2: creating replication slot ...
 demo2: creating node entry for local node ...
 demo3: creating replication slot ...
 demo3: creating node entry for local node ...
Creating base backup of the remote node...
63655/63655 kB (100%), 1/1 tablespace
Creating restore point on remote node ...
Bringing local node to the restore point ...

它就在那里。我假设这两个问题的原因相同。据我所知,本地节点(db2)上没有创建日志条目,但远程(db1)上存在以下内容

2016-10-12 22:38:43 UTC [20808-1] postgres@demo2 LOG:  logical decoding found consistent point at 0/5001F00
2016-10-12 22:38:43 UTC [20808-2] postgres@demo2 DETAIL:  There are no running transactions.
2016-10-12 22:38:43 UTC [20808-3] postgres@demo2 STATEMENT:  SELECT pg_create_logical_replication_slot('bdr_17163_6340711416785871202_2_17163__', 'bdr');
2016-10-12 22:38:43 UTC [20811-1] postgres@demo3 LOG:  logical decoding found consistent point at 0/5002090
2016-10-12 22:38:43 UTC [20811-2] postgres@demo3 DETAIL:  There are no running transactions.
2016-10-12 22:38:43 UTC [20811-3] postgres@demo3 STATEMENT:  SELECT pg_create_logical_replication_slot('bdr_17939_6340711416785871202_2_17939__', 'bdr');
2016-10-12 22:38:44 UTC [20812-1] postgres@demo3 LOG:  restore point "bdr_6340711416785871202" created at 0/50022A8
2016-10-12 22:38:44 UTC [20812-2] postgres@demo3 STATEMENT:  SELECT pg_create_restore_point('bdr_6340711416785871202')

有什么帮助吗?

【问题讨论】:

  • 在上游节点发出一个提交以确保 WAL 刷新超过还原点。顺便说一句,如果您通常不知道您在使用 Pg 做什么,则 BDR 不是一个可以使用的工具。您需要了解它施加的限制以及有效使用它所需的应用程序更改。请详细阅读文档。它不适合 PostgreSQL 的新用户。也许你应该只使用你所知道的?或者至少使用简单的 PostgreSQL 主动/备用复制和 repmgr?
  • 同时检查你运行bdr_init_copy的目录中的bdr_init_copy_postgres.log

标签: postgresql replication postgresql-9.4 postgresql-bdr


【解决方案1】:

好的,刚刚遇到这个问题,其他论坛都没有任何帮助。他们中的一些人甚至说新节点可以将其状态报告为“o”,而其他节点将新服务器状态报告为“i”,因为“这只是一个错误,很好”。不行。新服务器可以接收复制更新,但新服务器上无法进行主要更新。解决此问题的关键是加快您要加入的服务器(不是新服务器)上的日志记录。在新的服务器日志上,您可能会看到类似:08006: could not receive data from client: Connection reset by peer,这不是很有帮助,并且会让您检查防火墙等。真正的资金来自源服务器日志,当他们有类似以下内容的日志时: no free replication state could be found for 11, increase max_replication_slots 可能发生的情况是您的集群中有太多服务器无法使用默认设置,或者更有可能是旧主机遗留了一些垃圾。

您需要清理现有集群中的每台服务器上的内容(注意!)。首先获取现有集群上的事物列表:

select * from bdr.bdr_nodes order by node_sysid;

然后,检查以下内容:

select conn_sysid,conn_dboid from bdr.bdr_connections order by conn_sysid;

.. 如果您看到旧条目(不包含第一个查询中的 node_sysid),则删除 例如。 delete from bdr.bdr_connections where conn_sysid='<from-first-query>';

select * from pg_replication_slots order by slot_name;

.. 如果您看到不包含活动 sysid 的旧条目,则删除 .. 注意,使用该功能,不要执行“删除” 例如。 select pg_drop_replication_slot('bdr_17213_6574566740899221664_1_17213__');

select * from pg_replication_identifier order by riname;

.. 如果您看到不包含活动 sysid 的旧条目,则删除 .. 注意,使用函数,不要做“删除”

select pg_replication_identifier_drop('bdr_6443767151306784833_1_17210_17213_');

运气好的话,在每个节点上完成此操作后,您会看到新服务器的 BDR 状态变为“r”。当您清理每台主机时,您应该注意到日志“08006:无法从客户端接收数据:对等方重置连接”,匹配您刚刚清理的服务器的 conn-sysid,停止发生。祝你好运

【讨论】:

    猜你喜欢
    • 2017-06-29
    • 2011-05-07
    • 2021-04-23
    • 1970-01-01
    • 1970-01-01
    • 2019-04-29
    • 2019-08-01
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多