【问题标题】:Create histogram for grouped column为分组列创建直方图
【发布时间】:2018-11-26 07:27:09
【问题描述】:

我对 Python 完全陌生。如何创建一个包含一行和三列的图,在每列中我绘制一个直方图?数据来自这个DataFrame:

 import pandas as pd
import matplotlib as plt
d = {'col1': ['A','A','A','A','A','A','B','B','B','B','B','B','C','C','C','C','C','C'], 
     'col2': [3, 4, 3, 4, 6, 7, 8, 9, 3, 2, 3, 4, 5, 3, 4, 1, 2, 6 ]}
df = pd.DataFrame(data=d)

在 DataFrame 中,我们有三个组(A、B、C),但我可以有 N 个组,但我仍然想要一个包含一行的图表,每列是每个组的直方图。 谢谢!

【问题讨论】:

  • 直方图是什么意思? col1中'B'的计数,还是'B'对应的col2值的总和?
  • 这是来自 matplotlib 的 hist()。 matplotlib.org/1.2.1/examples/pylab_examples/…
  • 您的问题可以有不同的解释。我知道不需要一行子图,直方图确实应该在同一个图中。这是正确的吗?
  • 好的,抱歉。我希望每个组都有一个图,并且每个图应该彼此相邻,例如:A B C。下面的示例非常完美,但是我不需要将图彼此相邻,而是需要彼此相邻的图。谢谢

标签: python pandas matplotlib plot histogram


【解决方案1】:

您可以旋转数据框并链接 plot 命令以生成图形。

import pandas as pd
import matplotlib.pyplot as plt

d = {'Category': ['A','A','A','A','A','A','B','B','B','B','B','B','C','C','C','C','C','C'], 
     'Values': [3, 4, 3, 4, 6, 7, 8, 9, 3, 2, 3, 4, 5, 3, 4, 1, 2, 2 ]}
df = pd.DataFrame(d)

df.pivot(columns='Category', values='Values').plot(kind='hist', subplots=True, rwidth=0.9, align='mid')

编辑:您可以使用下面的代码将所有子图绘制在一行中。但是,对于三个以上的类别,地块开始看起来非常紧凑。

df2 = df.pivot(columns='Category', values='Values')
color = ['blue', 'green', 'red']
idx = np.arange(1, 4)
plt.subplots(1, 3)
for i, col, colour in zip(idx, df2.columns, color):
    plt.subplot(1, 3, i)
    df2.loc[:, col].plot.hist(label=col, color=colour, range=(df['Values'].min(), df['Values'].max()), bins=11)
    plt.yticks(np.arange(3))
    plt.legend()

【讨论】:

  • 我要找的就是这个。有没有办法让图在同一行?而不是让他们一个在另一个之下,让他们一个挨着另一个像:A B C
  • 我已编辑我的答案以创建具有所需格式的图。我的观点是,这不值得努力。如果您拥有三个以上的类别,则垂直堆叠的子图效果会更好。
  • 谢谢!我收到以下错误:AttributeError: module 'matplotlib' has no attribute 'subplots'
  • 我认为这是因为我输入了 import matplotlib as plt 而不是 import matplotlib.pyplot as plt。你能用修改后的导入命令试试吗?
【解决方案2】:

您可以创建一行子图并用直方图填充每个子图:

import pandas as pd
from matplotlib import pyplot as plt
from matplotlib.ticker import FormatStrFormatter

#define toy dataset
d = {'col1': ['A','A','A','A','A','A','B','B','B','B','B','B','C','C','C','C','C','C'], 
     'col2': [3, 4, 3, 4, 6, 7, 8, 9, 3, 2, 3, 4, 5, 3, 4, 1, 2, 6 ]}
df = pd.DataFrame(data=d)

#number of bins for histogram
binnr = 4
#group data in dataframe
g = df.groupby("col1")
#create subplots according to unique elements in col1, same x and y scale for better comparison
fig, axes = plt.subplots(1, len(g), sharex = True, sharey = True)
#just in case you will extend it to a 2D array later
axes = axes.flatten()

#minimum and maximum value of bins to have comparable axes for all histograms
binmin = df["col2"].min()
binmax = df["col2"].max()

#fill each subplot with histogram
for i, (cat, group) in enumerate(g): 
    axes[i].set_title("graph {} showing {}".format(i, cat))
    _counts, binlimits, _patches = axes[i].hist(group["col2"], bins = binnr, range = (binmin, binmax))

#move ticks to label the bin borders
axes[0].set_xticks(binlimits)
#prevent excessively long tick labels
axes[0].xaxis.set_major_formatter(FormatStrFormatter('%0.1f'))
plt.tight_layout()
plt.show()

示例输出:

【讨论】:

    【解决方案3】:

    我认为这是您搜索的代码:

    import pandas as pd
    import matplotlib.pyplot as plt
    d = {'col1': ['A','A','A','A','A','A','B','B','B','B','B','B','C','C','C','C','C','C'], 
         'col2': [3, 4, 3, 4, 6, 7, 8, 9, 3, 2, 3, 4, 5, 3, 4, 1, 2, 6 ]}
    df = pd.DataFrame(data=d)
    
    keys = sorted(df['col1'].unique())
    
    vals = []
    for k in keys:
        vals.append(sum(df.loc[df['col1'] == k]['col2']))
    
    print(vals)
    
    plt.bar(keys, vals)
    plt.show()
    

    这是您在此示例中得到的:

    问我是否需要解释(或只是谷歌它☻)。

    【讨论】:

    • 我对上面 KRKirov 的示例代码很感兴趣,但不要让图彼此下方,而是让图彼此相邻,谢谢。
    猜你喜欢
    • 2015-02-10
    • 2016-04-25
    • 1970-01-01
    • 1970-01-01
    • 2021-08-26
    • 2021-04-06
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多