【问题标题】:Search for regex within dict considering TimeInterval dict's key考虑到 TimeInterval 字典的键,在字典中搜索正则表达式
【发布时间】:2023-01-05 18:02:20
【问题描述】:

我有以下数据结构 - 字典列表。每个字典包含:{ip:x.x.x.x, timestamp: , message: "yyyyy"}:

list1 =[
{'ip': '11.22.33.44', 'timestamp': 1665480231699, 'message': '{"body": "Idle time larger than time period. retry:0"}', 'ingestionTime': 1665480263198},
{'ip': '11.22.33.42', 'timestamp': 1665480231698, 'message': '{"body": "Idle time larger than time period. retry:5"}', 'ingestionTime': 1665480263198}, 
{'ip': '11.22.33.44', 'timestamp': 1665480231698, 'message': '{"body": "Idle time larger than time period. retry:0"}', 'ingestionTime': 1665480263198}
]

此外,我有一个正则表达式列表(whitelist_metadata),我想在上面的dicts msg中搜索(MetricMsg),并检查(根据时间戳)它是否在时间间隔内出现X次(对于我们的示例1分钟) - 验证应该是单ip。

whitelist_metadata = [
  {
    'LogLevel': 'WARNING',
    'SpecificVersion': 'None',
    'TimeInterval(Min)': 1,
    'MetricMsg': 'DDR: XXXX count got lost',
    'AllowedOccurrenceInTimeInterval': 0   --> this means that we are allowing this msg always 
  },
  {
    'LogLevel': 'WARNING',
    'SpecificVersion': 'None',
    'TimeInterval(Min)': 1,
    'MetricMsg': 'Idle time larger than XXX time. retry: \\d ',     --> please notice it's a regex 
    'AllowedOccurrenceInTimeInterval': 5  --> this means that are allowing this msg only if happened not more than 5 times within 1min.
  }
]

我的本能想法是:

  1. 在每个 ip 的消息值上运行以搜索单个正则表达式匹配项(将在循环中运行,因为我们有多个正则表达式要搜索)。
  2. 一旦找到消息 - 保存它的时间戳并检查之前时间戳之间的差异......(猜测有 pandas 技巧可以更好地支持时间间隔检查,看到这个我还没有使用过:https://www.geeksforgeeks.org/how-to-group-data-by-time-intervals-in-python-pandas/)
    • 如果它在允许的 TimeInterval 内且 <= AllowedOccurrenceInTimeInterval - 从所有服务器 ip 时间戳消息列表中弹出它。
    • 否则 - 将其留在消息列表中

    我开始这样编码:

     import pandas as pd
     df = pd.DataFrame(list1)
     df['timestamp'] = pd.to_datetime(df['timestamp'], unit="ms")
     group_per_ip = df.sort_values('timestamp').groupby("ip")
     # for ip in group_per_ip.groups.keys():
     #   single_ip = group_per_ip.get_group(ip)
     single_ip  =  group_per_ip.get_group('11.22.33.44')
    

    现在我试图弄清楚如何在其上运行 pandas rolling("5m") 函数,但它一直抛出相同的错误: ValueError('window must be an integer',)

    我试着关注:Python, Pandas ; ValueError('window must be an integer',) 但没有帮助

    有人可以帮助我找到一种使用熊猫或其他处理此类 TimeInterval 问题的良好性能建议来实现它的方法吗?

【问题讨论】:

    标签: python pandas pivot


    【解决方案1】:

    分享杆子,我从那里继续..

    import pandas as pd
    import re
    import json
    
    list1 = [
        {'ip': '11.22.33.44', 'timestamp': 1665480231699,
            'message': '{"body": "Idle time larger than time period. retry:0"}', 'ingestionTime': 1665480263198},
        {'ip': '11.22.33.42', 'timestamp': 1665480231698,
            'message': '{"body": "Idle time larger than time period. retry:5"}', 'ingestionTime': 1665480263198},
        {'ip': '11.22.33.44', 'timestamp': 1665480231698,
            'message': '{"body": "Idle time larger than time period. retry:0"}', 'ingestionTime': 1665480263198}
    ]
    
    df = pd.DataFrame(list1)
    df['timestamp'] = pd.to_datetime(df['timestamp'], unit="ms")
    group_per_ip = df.sort_values('timestamp').groupby("ip")
    # for ip in group_per_ip.groups.keys():
    #   single_ip = group_per_ip.get_group(ip)
    single_ip = group_per_ip.get_group('22.132.44.253')
    
    # %%
    single_ip_index = single_ip.set_index('timestamp')
    single_ip_temp = single_ip_index['message'].apply(func=lambda x: re.match(
        r"Idle time larger than time period. retry:d+", json.loads(x)['body']) is not None)
    single_ip_temp.rolling('2T').count() > 1.5
    

    【讨论】:

      猜你喜欢
      • 2022-10-13
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 2013-04-14
      • 2015-05-26
      • 2016-06-01
      • 2018-11-21
      • 2012-06-11
      相关资源
      最近更新 更多