【问题标题】:Python multiprocessing creates incorrect pidsPython 多处理创建不正确的 pid
【发布时间】:2012-08-09 13:45:29
【问题描述】:

我已经为此工作了一段时间,但似乎无法弄清楚。我把它缩小到我的代码和操作系统(Linux 上的 Python 2.7.3)不同意应该运行哪些进程的情况。发生这种情况时,我的代码将永远挂起,但不会引发异常。有时代码会正常运行几个小时,有时只运行几分钟,我不知道为什么。这表现如下。感谢您的关注,我真的很喜欢这里(双关语)。

代码输出:

创建离散字符矩阵

running PoolWorker_82 (72 triplets), pid 25777, ppid 24892
running PoolWorker_83 (72 triplets), pid 25778, ppid 24892
running PoolWorker_84 (72 triplets), pid 25779, ppid 24892
running PoolWorker_85 (72 triplets), pid 25780, ppid 24892
running PoolWorker_86 (72 triplets), pid 25781, ppid 24892
running PoolWorker_87 (72 triplets), pid 25782, ppid 24892
running PoolWorker_88 (72 triplets), pid 25783, ppid 24892
running PoolWorker_89 (90 triplets), pid 25784, ppid 24892

ps aux 的输出...

1000     24892  2.0  0.9 559948 151088 pts/0   Sl+  09:14   0:16 p runsimulation.py
1000     25776  0.0  0.8 559932 138320 pts/0   S+   09:19   0:00 p runsimulation.py
1000     26015  0.0  0.8 559948 138140 pts/0   S+   09:22   0:00 p runsimulation.py
1000     26021  0.0  0.8 559948 138140 pts/0   S+   09:22   0:00 p runsimulation.py
1000     26023  0.0  0.8 559948 138140 pts/0   S+   09:22   0:00 p runsimulation.py
1000     26025  0.0  0.8 559948 138140 pts/0   S+   09:22   0:00 p runsimulation.py
1000     26027  0.0  0.8 559948 138140 pts/0   S+   09:22   0:00 p runsimulation.py
1000     26029  0.0  0.8 559948 138140 pts/0   S+   09:22   0:00 p runsimulation.py
1000     26031  0.0  0.8 559948 138140 pts/0   S+   09:22   0:00 p runsimulation.py
1000     26036  0.0  0.8 559948 138140 pts/0   S+   09:22   0:00 p runsimulation.py

您可以看到父进程 24982 在那里,但工人的 pid 不存在。通常情况下,这些会匹配,我可以看到工作人员在工作时 CPU 使用率达到 100%,然后在迭代完成后它们都消失了。当它失败时,我得到 pid 不匹配和使用 0.0% CPU 的进程(第 3 列)。

我的代码的相关部分如下(以它们被调用的相反顺序):

使用 rpy2 调用的函数的 R 设置:

def create_R(dir):
    """
    creates the r environment
    @param dir: the directory for the output files
    """
    r = robjects.r
    importr("phangorn")
    importr("picante")
    importr("MASS")
    importr("vegan")
    r("options(expressions=500000)")
    robjects.globalenv['outfile'] = os.path.abspath(os.path.join(dir, "trees.pdf"))
    r('pdf(file=outfile, onefile=T)')
    r("par(mfrow=c(2,3))")

    r("""
        generate_triplet = function(bits) {
        triplet = replicate(bits, rTraitDisc(tree, model="ER", k=2,states=0:1))
        triplet = t(apply(triplet, 1, as.numeric))
        sums = rowSums(triplet)
        if (length(which(sums==0)) > 0 && length(which(sums==3)) == 1) {
            return(triplet)
        }
        return(generate_triplet(bits))
        }
    """)

    r("""
        get_valid_triplets = function(numsamples, needed, bits) {
            tryCatch({
                m = generate_triplet(bits)
                while (ncol(m) < needed) {
                    m = cbind(m, generate_triplet(bits))
                }
            return(m)
            }, error = function(e){print(message(e))}, warning = function(e){print(message(e))})
        }
    """)

worker 内部调用的函数:

def __get_valid_triplets(num_samples, num_triplets, bits, q):
    r = robjects.r
    name = current_process().name.replace("-", "_")
    timer = stopwatch.Timer()
    log("\trunning %s (%d triplets), pid %d, ppid %d" % (name, num_triplets, current_process().pid, os.getppid()),
        log_file)
    r('%s = get_valid_triplets(%d, %d, %d)' % (name, num_samples, num_triplets, bits))
    q.put((name, r[name]))
    timer.stop()
    log("\t%s complete (%s)" % (name, str(timer)), log_file)

设置池并使用 apply_async 调度工作人员的函数。唤醒者写入一个托管队列,该队列是池加入后的进程:

def __generate_candidate_discrete_matrix(num_cols, num_samples, sample_tree, bits, usable_cols):
    assert isinstance(sample_tree, dendropy.Tree)
    print "Creating discrete character matrix"
    r = robjects.r
    newick = sample_tree.as_newick_string()
    num_samples = len(sample_tree.leaf_nodes())
    robjects.globalenv['numcols'] = usable_cols
    robjects.globalenv['newick'] = newick + ";"
    r("tree = read.tree(text=newick)")
    r('m = matrix(nrow=length(tree$tip.label))') #create empty matrix
    r('m = m[,-1]') #drop the first NA column
    num_procs = mp.cpu_count()
    args = []
    div, mod = divmod(usable_cols, num_procs)
    [args.append(div) for i in range(num_procs)]
    args[-1] += mod
    for i, elem in enumerate(args):
        div, mod = divmod(elem, bits)
        args[-1] += mod
        args[i] -= mod
    manager = Manager()
    pool = Pool(processes=num_procs, maxtasksperchild=1)
    q = manager.Queue(maxsize=num_procs)
    for arg in args:
        pool.apply_async(__get_valid_triplets, (num_samples, arg, bits, q))
    pool.close()
    pool.join()

    while not q.empty():
        name, data = q.get()
        robjects.globalenv[name] = data
        r('m = cbind(m, %s)' % name)

    r('m = m[,1:%d]' % usable_cols)
    r('m = m[order(rownames(m)),]') # consistently order the rows 
    r('m = t(apply(m, 1, as.numeric))') # convert all factors given by rTraitDisc to numeric
    a = r['m']
    n = r('rownames(m)')
    return a, n

最后,调用的第一个函数生成候选矩阵,确保它是一个有效的,如果不是,它会再次尝试一个新的矩阵。如果有效,则在 R session 中存储一些东西并返回数据

def create_discrete_matrix(num_cols, num_samples, sample_tree, bits):
    """
    Creates a discrete char matrix from a tree
    @param num_cols: number of columns to create
    @param sample_tree: the tree
    @return: a r object of the matrix, and a list of the row names
    @rtype: tuple(robjects.Matrix, list)
    """
    r = robjects.r
    usable_cols = find_usable_length(num_cols, bits)
    a, n = __generate_candidate_discrete_matrix(num_cols, num_samples, sample_tree, bits, usable_cols)
    assert isinstance(a, robjects.Matrix)
    assert a.ncol == usable_cols

    paralin_matrix, valid = __create_paralin_matrix(a)
    if valid is False:
        sample_tree = create_tree(num_samples, type = "S")
        return create_discrete_matrix(num_cols, num_samples, sample_tree, bits)
    else:
        robjects.globalenv['paralin_matrix'] = paralin_matrix
        r('rownames(paralin_matrix) = rownames(m)')
        r('paralin_dist = as.dist(paralin_matrix, diag=T, upper=T)')
        r("paralinear_cluster = hclust(paralin_dist, method='average')")
    return sample_tree, a, n

【问题讨论】:

  • 好吧,您很可能在这里遇到了 python 多处理中的错误。尝试进一步缩小生成错误的代码,并在 bugs.python.org 上的 Python 错误跟踪器上打开一个错误
  • 谢谢,jsbueno。已经完成了!错误报告在这里:bugs.python.org/issue15603。该错误一定发生在工作人员的某个地方,但所有追踪它的努力都失败了。我倾向于只在我走到尽头的时候才在这里发帖,就在我把笔记本电脑扔出窗外之前。
  • 这真的是重现问题的最小例子吗?你是不是一开始就说,让工人什么都不做,只导入 rpy2 看看问题是否已经存在?
  • 这绝对不是问题的最小例子。有一次我退后一步写了一个,不小心在worker的print语句中出错了。当没有像应该那样发送回溯时,我意识到它没有办法传播回主线程​​(在那里有一个尝试)。因此,在我的代码中,我在 __get_valid_triplets 中明确添加了一个 try/except ,它将在池中的每个工作进程中起作用。手指交叉,现在等待追溯......

标签: python multiprocessing rpy2


【解决方案1】:

这似乎已通过服务器重启 (FML) 得到修复。但是,获得了有效信息。将 worker 提交到池时,请确保在 worker 本身中捕获异常,而不是在调用 pool.apply_async 的方法中捕获它们。

def __get_valid_triplets(num_samples, num_triplets, bits, q):
    try:
        r = robjects.r
        name = current_process().name.replace("-", "_")
        timer = stopwatch.Timer()
        log("\trunning %s (%d triplets), pid %d, ppid %d" % (name, num_triplets, current_process().pid, os.getppid()),
            log_file)
        r('%s = get_valid_triplets(%d, %d, %d)' % (name, num_samples, num_triplets, bits))
        q.put((name, r[name]))
        timer.stop()
        log("\t%s complete (%s)" % (name, str(timer)), log_file)
    except Exception, e:
        q.put("DEATH")
        traceback.print_exc()

【讨论】:

    猜你喜欢
    • 2013-02-25
    • 2016-06-05
    • 2010-10-15
    • 1970-01-01
    • 2021-07-31
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多