【问题标题】:Monitor a Pacemaker Cluster with ocf:pacemaker:ClusterMon and/or external-agent使用 ocf:pacemaker:ClusterMon 和/或 external-agent 监控 Pacemaker 集群
【发布时间】:2014-11-21 09:57:14
【问题描述】:

我正在尝试通过外部代理配置 Pacemaker 集群事件通知,以便在发生故障转移切换时接收通知。
我搜索了以下链接

https://access.redhat.com/documentation/en-US/Red_Hat_Enterprise_Linux/6/html/Configuring_the_Red_Hat_High_Availability_Add-On_with_Pacemaker/s1-eventnotification-HAAR.html

http://floriancrouzat.net/2013/01/monitor-a-pacemaker-cluster-with-ocfpacemakerclustermon-andor-external-agent/

但不了解实际如何做到这一点。
能否请您逐步解释一下。

谢谢你,
兰詹。

【问题讨论】:

    标签: linux redhat


    【解决方案1】:

    RedHat 文档很简洁,但 Florian 的博客条目非常详细,最后的参考资料很有帮助。

    所问的问题有点模糊,所以我正在回答我认为你在问的问题。

    简要地总结一下 Florian 的帖子,ClusterMon 是一个资源代理 (ocf:pacemaker:ClusterMon),它在幕后运行 crm_mon。

    我的 (SLES 11 SP3) 资源文档说:

    # crm ra info ocf:pacemaker:ClusterMon
    Runs crm_mon in the background, recording the cluster status to an HTML file (ocf:pacemaker:ClusterMon)
    
    This is a ClusterMon Resource Agent.
    It outputs current cluster status to the html.
    
    Parameters (* denotes required, [] the default):
    
    user (string, [root]): 
        The user we want to run crm_mon as
    
    update (integer, [15]): Update interval
        How frequently should we update the cluster status
    
    extra_options (string): Extra options
        Additional options to pass to crm_mon.  Eg. -n -r
    
    pidfile (string, [/tmp/ClusterMon_undef.pid]): PID file
        PID file location to ensure only one instance is running
    
    htmlfile (string, [/tmp/ClusterMon_undef.html]): HTML output
        Location to write HTML output to.
    
    Operations' defaults (advisory minimum):
    
        start         timeout=20
        stop          timeout=20
        monitor       timeout=20 interval=10
    

    但是,真正的力量是extra_options,因为这使您可以让资源代理告诉crm_mon 如何处理结果。具体来说,extra_options 作为crm_mon 的命令行选项逐字传递。

    正如弗洛里安所提到的,最近的 crm_mon(实际工作)没有内置 SMTP(电子邮件)或 SNMP 支持。但是,它仍然支持外部代理(通过 -E 开关)。

    因此,要了解 extra_options 的作用,您应该咨询man crm_mon。

    从您链接到的 RedHat 文档中,-T pacemaker@example.com -F pacemaker@nodeX.example.com -P PACEMAKER -H mail.example.com 的第一个“extra_options”值告诉 crm_mon 发送电子邮件至 pacemaker@example.com,来自 pacemaker@nodeX.example.com,主题前缀为 PACEMAKER,通过邮件主机(smtp 服务器)mail.example.com。

    您引用的 RedHat 文档中的第二个“extra_options”示例具有值 -S snmphost.example.com -C public,它告诉 crm_mon 使用名为 public 的社区将 SNMP 陷阱发送到 snmphost.example.com。

    第三个“extra_options”示例的值为-E /usr/local/bin/example.sh -e 192.168.12.1。这告诉crm_mon 运行外部程序/usr/local/bin/example.sh,它还指定了“外部收件人”,它实际上只是被扔进了一个环境变量CRM_notify_recipient,它在生成脚本之前被导出。

    运行外部代理时,crm_mon 调用为每个 集群事件(包括成功的监控操作!)提供的脚本。这个脚本继承了一堆环境变量,告诉你发生了什么。

    发件人:http://clusterlabs.org/doc/en-US/Pacemaker/1.1/html/Pacemaker_Explained/s-notification-external.html 设置的环境变量是:

    CRM_notify_recipient    The static external-recipient from the resource definition.
    CRM_notify_node The node on which the status change happened.
    CRM_notify_rsc  The name of the resource that changed the status.
    CRM_notify_task The operation that caused the status change.
    CRM_notify_desc The textual output relevant error code of the operation (if any) that caused the status change.
    CRM_notify_rc   The return code of the operation.
    CRM_notify_target_rc    The expected return code of the operation.
    CRM_notify_status   The numerical representation of the status of the operation.
    

    脚本的工作是使用这些环境变量并对它们做一些合理的事情。什么是“合理”取决于您的环境。

    Florian 博客中的 SNMP 陷阱示例假设您熟悉 SNMP 陷阱。如果不是,那么这是一个完全不同的问题,超出了资源代理的范围。

    带有 SNMP 陷阱的示例提供了一个很好的条件语句来识别事件是不成功的监控事件或不是监控事件的事件。

    监视脚本的脚手架可以根据可用信息执行任何操作,这实际上是 Florian 博客文章中引用的 snmp 陷阱 shell 脚本的精简版本。它看起来像:

    #!/bin/bash
    
    # if [[ unsuccessful monitor operation ]] or [[ not monitor op ]]
    if [[ ${CRM_notify_rc} != 0 && ${CRM_notify_task} == "monitor" ]] || \
       [[ ${CRM_notify_task} != "monitor" ]] ; then
    
        # Do whatever you want with the information available in the
        # environment variables mentioned above that will do something
        # meaningful for you.
    
        # EG: Fire off an email attempting to be human readable
        # SUBJ="${CRM_notify_task} ${CRM_notify_desc} for ${CRM_notify_rsc} "
        # SUBJ="$SUBJ on ${CRM_notify_node}"
        # MSG="The ${CRM_notify_task} operation for ${CRM_notify_rsc} on "
        # MSG="$MSG ${CRM_notify_node} exited with status ${CRM_notify_rc} "
        # MSG="$MSG (${CRM_notify_desc}) and we expected ${CRM_notify_target_rc}"
        # echo "$MSG" | mail -s "$SUBJ" you@host.com
    
    
    fi
    exit 0 
    

    但是,如果您遵循 Florian 的建议并克隆资源,则脚本将在每个节点上运行。对于 SNMP 陷阱,这非常好。但是,如果您正在执行诸如从脚本发送电子邮件之类的操作,您可能不想实际克隆它。

    【讨论】:

      【解决方案2】:

      两个节点:

      cat << 'EOL'>/usr/local/bin/crm_e-mail.sh
      #!/bin/bash
      echo "Please check your installation @ http://domain.com.com:2224 & http://domain.com.com/clustermon.html" | mail -s "Cluster Change Detected" sysalert@domain.com
      EOL
      chmod 700 /usr/local/bin/crm_e-mail.sh
      chown root.root /usr/local/bin/crm_e-mail.sh
      

      一个节点:

      pcs resource create ClusterMon-SMTP ClusterMon user=root \
      update=10 extra_options="-E /usr/local/bin/crm_e-mail.sh --watch-fencing" \
      pidfile=/var/run/crm_mon-smtp.pid clone
      

      【讨论】:

      猜你喜欢
      • 2021-10-31
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 2017-02-18
      • 2016-07-19
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      相关资源
      最近更新 更多