【问题标题】:Connection lock using infinispan and hibernate on distributed cache l2在分布式缓存 l2 上使用 infinispan 和休眠进行连接锁定
【发布时间】:2018-12-19 23:10:54
【问题描述】:

我使用 infinispan 作为分布式休眠缓存 L2,使用 AWS 上的 jgroups 进行配置。 但是,我在以下情况下面临重负载问题:

  • 最初有 2 个 EC2 实例可用
  • 负载急剧增加
  • 负载平衡启动 4 个 EC2 实例
  • 负载部分减少
  • 负载均衡减少 2 个 EC2 实例
  • 剩余的所有实例都开始保持数据库连接(Hikari 池),直到没有可用时,执行所有后续请求都会返回错误,因为等待空闲的超时。

其余实例尝试与旧实例通信但没有得到响应,在等待响应时保持连接。

所有实体都在使用 READ_WRITE 策略。

Infinispan 配置: org/infinispan/hibernate/cache/commons/builder/infinispan-configs.xml

region.factory_class: org.infinispan.hibernate.cache.commons.InfinispanRegionFactory

以下 Jgroups 配置编辑自: org/infinispan/infinispan-core/9.2.0.Final/infinispan-core-9.2.0.Final.jar/default-configs/default-jgroups -tcp.xml

    <config xmlns="urn:org:jgroups"
        xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance"
        xsi:schemaLocation="urn:org:jgroups http://www.jgroups.org/schema/jgroups-4.0.xsd">
   <TCP bind_port="7800"
        enable_diagnostics="false"
        thread_naming_pattern="pl"
        send_buf_size="640k"
        sock_conn_timeout="300"
        bundler_type="no-bundler"

        thread_pool.min_threads="${jgroups.thread_pool.min_threads:50}"
        thread_pool.max_threads="${jgroups.thread_pool.max_threads:500}"
        thread_pool.keep_alive_time="30000"
   />
   <AWS_ELB_PING
           region="sa-east-1"
           load_balancers_names="elb-name"
   />
   <MERGE3 min_interval="5000"
           max_interval="30000"
   />
   <FD_SOCK />
   <FD_ALL timeout="9000"
           interval="3000"
           timeout_check_interval="1000"
   />
   <VERIFY_SUSPECT timeout="5000" />
   <pbcast.NAKACK2 use_mcast_xmit="false"
                   xmit_interval="100"
                   xmit_table_num_rows="50"
                   xmit_table_msgs_per_row="1024"
                   xmit_table_max_compaction_time="30000"
                   resend_last_seqno="true"
   />
   <UNICAST3 xmit_interval="100"
             xmit_table_num_rows="50"
             xmit_table_msgs_per_row="1024"
             xmit_table_max_compaction_time="30000"
             conn_expiry_timeout="0"
   />
   <pbcast.STABLE stability_delay="500"
                  desired_avg_gossip="5000"
                  max_bytes="1M"
   />
   <pbcast.GMS print_local_addr="false"
               install_view_locally_first="true"
               join_timeout="${jgroups.join_timeout:5000}"
   />
   <MFC max_credits="2m"
        min_threshold="0.40"
   />
   <FRAG3/>
</config>

AWS_ELB_PING:这个类是Discovery类的一个实现,使用AWS ELB api来发现所有可用的ip。

我从以下代码中删除了日志和一些样板代码:

public class AWS_ELB_PING extends Discovery {

    private static final String LIST_ELEMENT_SEPARATOR = ",";

    static {
        ClassConfigurator.addProtocol((short) 790, AWS_ELB_PING.class); // id must be unique
    }

    private String region;

    private String load_balancers_names;

    private int bind_port = 7800;

    private AmazonElasticLoadBalancing amazonELBClient;
    private AmazonEC2 amazonEC2Client;

    private List<String> getLoadBalancersNamesList() {
        return Arrays.asList(Optional.ofNullable(load_balancers_names).orElse("").split(LIST_ELEMENT_SEPARATOR));
    }

    @Override
    public void init() throws Exception {
        super.init();
        DefaultAWSCredentialsProviderChain awsCredentialsProviderChain = DefaultAWSCredentialsProviderChain.getInstance();
        amazonELBClient = AmazonElasticLoadBalancingClientBuilder.standard()
                .withRegion(region)
                .withCredentials(awsCredentialsProviderChain)
                .build();
        amazonEC2Client = AmazonEC2ClientBuilder.standard()
                .withRegion(region)
                .withCredentials(awsCredentialsProviderChain)
                .build();
    }

    @Override
    public void discoveryRequestReceived(final Address sender, final String logical_name,
                                         final PhysicalAddress physical_addr) {
        super.discoveryRequestReceived(sender, logical_name, physical_addr);
    }

    @Override
    public void findMembers(final List<Address> members, final boolean initialDiscovery, final Responses responses) {
        PhysicalAddress physicalAddress = null;
        PingData data = null;
        if (!use_ip_addrs || !initialDiscovery) {
            physicalAddress = (PhysicalAddress) super.down(new Event(Event.GET_PHYSICAL_ADDRESS, local_addr));
            data = new PingData(local_addr, false, NameCache.get(local_addr), physicalAddress);
            if (members != null && members.size() <= max_members_in_discovery_request) {
                data.mbrs(members);
            }
        }

        sendDiscoveryRequests(physicalAddress, data, initialDiscovery, getLoadBalancersInstances());
    }

    private Set<Instance> getLoadBalancersInstances() {
        final List<String> loadBalancerNames = getLoadBalancersNamesList();
        final List<LoadBalancerDescription> loadBalancerDescriptions = amazonELBClient
                .describeLoadBalancers(new DescribeLoadBalancersRequest().withLoadBalancerNames(loadBalancerNames))
                .getLoadBalancerDescriptions();

        checkLoadBalancersExists(loadBalancerNames, loadBalancerDescriptions);

        final List<String> instanceIds = loadBalancerDescriptions.stream()
                .flatMap(loadBalancer -> loadBalancer.getInstances().stream())
                .map(instance -> instance.getInstanceId())
                .collect(toList());

        return amazonEC2Client.describeInstances(new DescribeInstancesRequest().withInstanceIds(instanceIds))
                    .getReservations()
                    .stream()
                    .map(Reservation::getInstances)
                    .flatMap(List::stream)
                    .collect(Collectors.toSet());
    }

    private void checkLoadBalancersExists(final List<String> loadBalancerNames,
                                          final List<LoadBalancerDescription> loadBalancerDescriptions) {
        final Set<String> difference = Sets.difference(new HashSet<>(loadBalancerNames),
            loadBalancerDescriptions
                    .stream()
                    .map(LoadBalancerDescription::getLoadBalancerName)
                    .collect(Collectors.toSet()));
    }

    private PhysicalAddress toPhysicalAddress(final Instance instance) {
        try {
            return new IpAddress(instance.getPrivateIpAddress(), bind_port);
        } catch (final Exception e) {
            throw new RuntimeException(e);
        }
    }

    private void sendDiscoveryRequests(@Nullable final PhysicalAddress localAddress, @Nullable final PingData data,
                                       final boolean initialDiscovery, final Set<Instance> instances) {
        final PingHeader header = new PingHeader(PingHeader.GET_MBRS_REQ)
                    .clusterName(cluster_name)
                    .initialDiscovery(initialDiscovery);

        instances.stream()
                .map(this::toPhysicalAddress)
                .filter(physicalAddress -> !physicalAddress.equals(localAddress))
                .forEach(physicalAddress -> sendDiscoveryRequest(data, header, physicalAddress));
    }

    private void sendDiscoveryRequest(@Nullable final PingData data, final PingHeader header,
                      final PhysicalAddress destinationAddress) {
        final Message message = new Message(destinationAddress)
                .setFlag(Message.Flag.INTERNAL, Message.Flag.DONT_BUNDLE, Message.Flag.OOB)
                .putHeader(this.id, header);
        if (data != null) {
            message.setBuffer(marshal(data));
        }

        if (async_discovery_use_separate_thread_per_request) {
            timer.execute(() -> sendDiscoveryRequest(message), sends_can_block);
        } else {
            sendDiscoveryRequest(message);
        }
    }

    protected void sendDiscoveryRequest(final Message message) {
        try {
            super.down(message);
        } catch (final Throwable t) {
        }
    }

    @Override
    public boolean isDynamic() {
        return true;
    }

    @Override
    public void stop() {
        try {
            if (amazonEC2Client != null) {
                amazonEC2Client.shutdown();
            }
            if (amazonELBClient != null) {
                amazonELBClient.shutdown();
            }
        } catch (final Exception e) {
        } finally {
            super.stop();
        }
    }
}

有人遇到过这种问题吗?

【问题讨论】:

  • 在超时开始出现时或之前是否有一些线程转储?这样您就可以确切地看到线程在等待什么。

标签: java hibernate amazon-web-services infinispan jgroups


【解决方案1】:

你的问题是什么;你已经描述了你的设置,但不是你的问题...... 你有 AWS_ELP_PING 的参考吗?

【讨论】:

  • 缩减后剩余的机器保持连接,并且轮询用完空闲的,执行所有后续请求返回超时错误等待空闲连接。我将使用此信息更新问题并添加 AWS_ELP_PING 代码。
【解决方案2】:

代码对我来说看起来不错,尽管您可能想扩展现有的发现代码,例如TCPPING 或 FILE_PING,甚至 NATIVE_S3_PING。

“池耗尽”是什么意思?这是维护连接池的 AWS 客户端吗?或者你的意思是TCP中的线程池?后者有bindler_type=no-bundler;尝试删除它(然后将使用transfer-queue-bundler,它会创建消息批次而不是一个接一个地发送消息)。

如果你有一个耗尽的 TCP 线程池,那么获取堆栈跟踪会很有趣,看看现有线程在什么上被阻塞...

【讨论】:

  • 我已经尝试过使用 NATIVE_S3_PING 并且遇到了同样的问题这就是为什么我开发了这个新的,你建议的另外两个我还没有尝试过。关于连接池,忘记写“数据库”了。
  • 现在我很困惑:如果您的问题与耗尽的数据库连接池有关,这与 Infinispan 或 JGroups 有什么关系?
  • 我正在使用infinispan和jgroups来实现分布式缓存L2,只有开启缓存时才会出现这个问题。
猜你喜欢
  • 2012-03-09
  • 1970-01-01
  • 2021-10-03
  • 1970-01-01
  • 2015-11-07
  • 2013-09-07
  • 2014-01-21
  • 2021-02-18
  • 2023-03-03
相关资源
最近更新 更多