【问题标题】:Can I render only part of a scatter_matrix?我可以只渲染 scatter_matrix 的一部分吗?
【发布时间】:2017-03-26 04:22:09
【问题描述】:

在 Python 上,使用 Pandas 库,我尝试使用 scatter_matrix 生成 DataFrame 的散点图,如下所示:

scatter_matrix(df, alpha=0.5, figsize=(14,14), diagonal='kde')

该程序需要很长时间才能运行并最终崩溃,可能是因为有太多 (26) 列,并且生成的图像会太大。更不用说,我注意到我能够很好地渲染 13 个变量。这样,一种解决方案是生成 4 个图,生成的散布矩阵的每个象限一个,即范围[[0,0],[13,13]], [[13,0],[26,13]], [[0,13],[13,26]], [[13,13],[26,26]]。请注意,这些不是指源 DataFrame 上的范围,而是指我正在渲染的目标散布矩阵。有可能吗?

【问题讨论】:

  • 您是否尝试过在数据框的子集上调用 scatter_matrix? scatter_matrix(df[[columns_you_care_about]])?
  • @pault 问题是,假设我的 df 有列 ['a', 'b', 'c', 'd']。使用你的方法,我可以绘制两个象限,['a','b'] x ['a','b'] 和['c','d'] x ['c','d']。但是如何绘制剩余的 2 个象限 ['a','b'] x ['c','d'] 和 ['c','d'] x ['a','b']?

标签: python pandas matplotlib scipy


【解决方案1】:

四个象限将是:

mid_ind = len(df.index)//2
mid_col = len(df.columns)//2

df.iloc[:mid_ind,:mid_col]
df.iloc[mid_ind:,:mid_col]
df.iloc[mid_ind:,mid_col:]
df.iloc[:mid_ind,mid_col:]

【讨论】:

  • @MaiaVictor 是的,iloc 是你的新朋友。
  • 其实...重新考虑...它似乎没有做我想要的。查看this snippet 以了解不同之处。请告诉我,我不会发疯……这是不同的,对吧?
  • 此方法切分象限。如果您希望其他切片适应您在切片中使用的参数。如果您谈论 a b * a b 那么它不再是您在谈论的象限,而不是您的 OP 是关于什么
  • 我想你是对的,我已经编辑了 OP 以更好地表达我想要做的事情。现在正确吗?
  • 什么不清楚? scatter_plot 在我的数据集上生成一个 26x26 矩阵,我想将它分成四个 13x13 块……就是这样……
【解决方案2】:

我找不到任何官方方法,所以我修改了 scatter_matrix 实现以接收 2 个额外的参数,cols 和 rows,它们是带有您要比较的标签的数组:

"""
This is a modification of the scatter_matrix method;
it allows plotting sub-sections of the scatter plot
matrix. With the official method you can only plot
the entire matrix. This method allows selecting the
`cols` and `rows` you're interested in plotting.
Note that this wouldn't be possible even by calling
scatter_matrix on a subset of the dataframe, because
this wouldn't allow comparing different `cols`/`rows'.
"""

import numpy as np
import matplotlib.pyplot as plt
import pandas
import pandas.tools.plotting
from pandas.compat import range, lrange, lmap, map, zip, string_types

def scatter_matrix(frame, cols, rows, figsize=None, ax=None, grid=False,
                  diagonal='hist', marker='.', density_kwds=None,
                  hist_kwds=None, range_padding=0.05, **kwds):
    """
    Draw a matrix of scatter plots.
    Parameters
    ----------
    frame : DataFrame
    cols : [str]
      labels of the columns to be rendered
    rows:  [str]
      labels of the rows to be rendered
    alpha : float, optional
        amount of transparency applied
    figsize : (float,float), optional
        a tuple (width, height) in inches
    ax : Matplotlib axis object, optional
    grid : bool, optional
        setting this to True will show the grid
    diagonal : {'hist', 'kde'}
        pick between 'kde' and 'hist' for
        either Kernel Density Estimation or Histogram
        plot in the diagonal
    marker : str, optional
        Matplotlib marker type, default '.'
    hist_kwds : other plotting keyword arguments
        To be passed to hist function
    density_kwds : other plotting keyword arguments
        To be passed to kernel density estimate plot
    range_padding : float, optional
        relative extension of axis range in x and y
        with respect to (x_max - x_min) or (y_max - y_min),
        default 0.05
    kwds : other plotting keyword arguments
        To be passed to scatter function
    Examples
    --------
    >>> df = DataFrame(np.random.randn(1000, 4), columns=['A','B','C','D'])
    >>> scatter_matrix(df, alpha=0.2)
    """
    import matplotlib.pyplot as plt

    df = frame._get_numeric_data()
    w = len(cols)
    h = len(rows)
    naxes = w * h
    fig, axes = pandas.tools.plotting._subplots(naxes=naxes, figsize=figsize, ax=ax,
                          squeeze=False)

    # no gaps between subplots
    fig.subplots_adjust(wspace=0, hspace=0)

    #mask = pandas.tools.plotting.notnull(df)

    marker = pandas.tools.plotting._get_marker_compat(marker)

    hist_kwds = hist_kwds or {}
    density_kwds = density_kwds or {}

    # workaround because `c='b'` is hardcoded in matplotlibs scatter method
    #kwds.setdefault('c', plt.rcParams['patch.facecolor'])

    cols_boundaries_list = []
    for a in cols:
        values = df[a]#.values[mask[a].values]
        rmin_, rmax_ = np.min(values), np.max(values)
        rdelta_ext = (rmax_ - rmin_) * range_padding / 2.
        cols_boundaries_list.append((rmin_ - rdelta_ext, rmax_ + rdelta_ext))
    rows_boundaries_list = []
    for a in rows:
        values = df[a]#.values[mask[a].values]
        rmin_, rmax_ = np.min(values), np.max(values)
        rdelta_ext = (rmax_ - rmin_) * range_padding / 2.
        rows_boundaries_list.append((rmin_ - rdelta_ext, rmax_ + rdelta_ext))

    for i, a in zip(lrange(w), cols):
        for j, b in zip(lrange(h), rows):
            ax = axes[i, j]

            if cols[i] == rows[j]:
                values = df[a]#.values[mask[a].values]

                # Deal with the diagonal by drawing a histogram there.
                if diagonal == 'hist':
                    ax.hist(values, **hist_kwds)

                elif diagonal in ('kde', 'density'):
                    from scipy.stats import gaussian_kde
                    y = values
                    gkde = gaussian_kde(y)
                    ind = np.linspace(y.min(), y.max(), 1000)
                    ax.plot(ind, gkde.evaluate(ind), **density_kwds)

                ax.set_xlim(cols_boundaries_list[i])

            else:
                #common = (mask[a] & mask[b]).values
                ax.scatter(df[b], df[a], marker=marker, **kwds)

                ax.set_xlim(rows_boundaries_list[j])
                ax.set_ylim(cols_boundaries_list[i])

            ax.set_xlabel(b)
            ax.set_ylabel(a)

            if j != 0:
                ax.yaxis.set_visible(False)
            if i != w - 1:
                ax.xaxis.set_visible(False)

    # what is that for?
    #if len(df.columns) > 1:
        #lim1 = cols_boundaries_list[0]
        #locs = axes[0][1].yaxis.get_majorticklocs()
        #locs = locs[(lim1[0] <= locs) & (locs <= lim1[1])]
        #adj = (locs - lim1[0]) / (lim1[1] - lim1[0])

        #lim0 = axes[0][0].get_ylim()
        #adj = adj * (lim0[1] - lim0[0]) + lim0[0]
        #axes[0][0].yaxis.set_ticks(adj)

        #if np.all(locs == locs.astype(int)):
            ## if all ticks are int
            #locs = locs.astype(int)
        #axes[0][0].yaxis.set_ticklabels(locs)

    pandas.tools.plotting._set_ticks_props(axes, xlabelsize=8, xrot=90, ylabelsize=8, yrot=0)

    return axes

我很粗心地修改了该代码,因此它可能已损坏但满足了我的需求。

【讨论】:

    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2012-05-31
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2012-07-06
    相关资源
    最近更新 更多