【问题标题】:Concat a dataframe's value to a dictionary's matching value将数据框的值连接到字典的匹配值
【发布时间】:2019-08-05 18:54:00
【问题描述】:

与本帖相关的切线:Customize Bokeh Unemployment Example: Replacing Percentage Value

起始码: https://docs.bokeh.org/en/latest/docs/gallery/texas.html

from bokeh.io import show
from bokeh.models import LogColorMapper
from bokeh.palettes import Viridis6 as palette
from bokeh.plotting import figure

from bokeh.sampledata.us_counties import data as counties
counties = { code: county for code, county in counties.items() if county["state"] == "tx" }

csv 数据:

I have a dictionary of county names:
{(48, 1): {'name': 'Anderson',
  'detailed name': 'Anderson County, Texas',
  'state': 'tx'}
{(48, 3): {'name': 'Andrews',
  'detailed name': 'Andrews County, Texas',
  'state': 'tx'}

and a dataframe created from a csv file of percentage values:
 {'Anderson': 21.0,
 'Andrews': 28.0,
 'Angelina': 31.0,
 'Aransas': 24.0,
 'Archer': 11.0,
 'Armstrong': 53.0,
 'Atascosa': 27.0,
 'Austin': 30.0,
 'Bailey': 42.0,
 'Bandera': 0.0}

我正在尝试将数据框的百分比值与字典中的县名合并。

from bokeh.models import LogColorMapper
from bokeh.palettes import Viridis6 as palette
from bokeh.plotting import figure, show
from bokeh.sampledata.us_counties import data as counties
import csv
import pandas as pd

pharmacy_concentration = {}
with open('resources/unemployment.csv', mode = 'r') as infile:
    next(infile)
    reader = csv.reader(infile, delimiter = ',', quotechar = '"')
    for row in reader:
        name, concentration = row 
            pharmacy_concentration[name] = float(concentration)

counties = { code: county for code, county in counties.items() if county["state"] == "tx" }
counties = pd.concat(pharmacy_concentration[concentration], on='name', 
how='left', keys='concentration')

counties

我收到一个显示百分比值的键错误,但不知道为什么。

预期输出:

 counties
 {(48, 1): {'name': 'Anderson',
 'detailed name': 'Anderson County, Texas',
 'state': 'tx', 'concentration': 21}

【问题讨论】:

  • 请包括整个字典(如果它们不是太大的话)
  • 已添加。 csv 对于每个县名都有一个唯一的值,没有重复。
  • 我不是说图片,我是说那段代码,所以我可以自己复制运行@BoredPando
  • 抱歉,我可以在帖子中添加 csv 文件吗?我不知道该怎么做。 csv 是两列:名称和浓度。如果你能得到 1 行来连接它应该适用于其余的。因此,在 csv 文件中只有列标题:姓名、浓度,下一行:Anderson, 20。
  • 尽量不要添加文件,你能做到print(df.head(10)) 并将其复制并粘贴到你的帖子中吗?另外,您的字典在您的帖子中不完整。人们无法像这样帮助你。

标签: python pandas csv bokeh concat


【解决方案1】:

感谢@Tony

from bokeh.models import LogColorMapper
from bokeh.palettes import Viridis256 as palette
from bokeh.plotting import figure, show
from bokeh.sampledata.us_counties import data as counties
import csv

pharmacy_concentration = {}
with open('resources/unemployment.csv', mode = 'r') as infile:
    reader = [row for row in csv.reader(infile.read().splitlines())]
    for row in reader:
        try:
            county_name, concentration = row
            pharmacy_concentration[county_name] = float(concentration)
        except Exception, error:
            print error, row

counties = { code: county for code, county in counties.items() if county["state"] == 
"tx" }
county_xs = [county["lons"] for county in counties.values()]
county_ys = [county["lats"] for county in counties.values()]
county_names = [county['name'] for county in counties.values()]
# Below is the line of code I was missing to make it work
county_pharmacy_concentration_rates = [pharmacy_concentration[counties[county] 
['name']] for county in counties if counties[county]['name'] in 
pharmacy_concentration]

【讨论】:

    【解决方案2】:

    如果我理解正确,这就是你想要做的:

    首先,我们在两个数据框中获取您的字典:county_names 和 csv_data。 之后我将它们转换为正确的格式,但这对你来说可能不是必需的:

    county_names = pd.DataFrame({'(48, 1)': {'name': 'Anderson', 'detailed name': 'Anderson County, Texas', 'state': 'tx'}, 
                                 '(48, 3)': {'name': 'Andrews', 'detailed name': 'Andrews County, Texas', 'state': 'tx'}}).T.reset_index().rename({'index': 'County_ID'}, axis=1)
    
    print(county_names)
      County_ID           detailed name      name state
    0   (48, 1)  Anderson County, Texas  Anderson    tx
    1   (48, 3)   Andrews County, Texas   Andrews    tx
    
     d = {'Anderson': 21.0,
     'Andrews': 28.0,
     'Angelina': 31.0,
     'Aransas': 24.0,
     'Archer': 11.0,
     'Armstrong': 53.0,
     'Atascosa': 27.0,
     'Austin': 30.0,
     'Bailey': 42.0,
     'Bandera': 0.0}
    
    csv_data = pd.DataFrame(d, index=[0]).melt(var_name='name', value_name='concentration')
    print(csv_data)
            name  concentration
    0   Anderson           21.0
    1    Andrews           28.0
    2   Angelina           31.0
    3    Aransas           24.0
    4     Archer           11.0
    5  Armstrong           53.0
    6   Atascosa           27.0
    7     Austin           30.0
    8     Bailey           42.0
    9    Bandera            0.0
    

    现在我们可以合并name 列上的数据:

    df_final = pd.merge(county_names, csv_data, on='name')
    
    print(df_final)
      County_ID           detailed name      name state  concentration
    0   (48, 1)  Anderson County, Texas  Anderson    tx           21.0
    1   (48, 3)   Andrews County, Texas   Andrews    tx           28.0
    

    注意
    您可以使用 pandas 轻松阅读csv file,只需使用:

    pd.read_csv(infile, delimiter = ',', quotechar = '"')
    

    【讨论】:

      猜你喜欢
      • 1970-01-01
      • 2021-05-26
      • 2020-04-09
      • 2020-09-10
      • 2018-02-10
      • 2016-07-21
      • 2022-01-11
      • 2020-08-03
      • 2022-11-20
      相关资源
      最近更新 更多