【问题标题】:How to create a sub data-frame from a given pandas data-frame?如何从给定的熊猫数据框创建子数据框?
【发布时间】:2019-02-08 12:10:37
【问题描述】:

我编写了一个从给定数据集读取数据并将整个 txt 文件转换为 pandas 数据帧的代码(经过一些预处理)

  • 纬度代表行并显示在列表中。
  • 经度表示列,并显示在单独的列表中。

现在,我想从我创建的原始数据框创建一个更小的数据框(以便更容易理解和解释数据)并执行计算。为此,我通过跳过每 10 个元素创建了一个大小为 18 的较小列。这工作得很好。让我们将此新列称为 new_column。

现在,我想要对每一行进行迭代,并针对第 k 行和新列 j 的每个值,将其添加到新矩阵或数据框。
例如。如果第 10 行和 new_column 12 的值为“x”,我想将这个“x”添加到相同的位置,但在新的数据框(或矩阵)中。

我已经编写了以下代码,但我不知道如何执行让我执行上述操作的部分。

import matplotlib.pyplot as plt
import pandas as pd
import numpy as np
from scipy import interpolate
# open the file for reading
dataset = open("Aug-2016-potential-temperature-180x188.txt", "r+")

# read the file linewise
buffer = dataset.readlines()

# pre-process the data to get the columns
column = buffer[8]
column = column[3 : -1]

# get the longitudes as features
features = column.split("\t")

# convert the features to float data-type
longitude = []

for i in features:
    if "W" in features:
        longitude.append(-float(i[:-1]))   # append -ve sign if "W", drop the "W" symbol
    else:
        longitude.append(float(i[:-1]))    # append +ve sign if "E", drop the "E" symbol

# append the longitude as columns to the dataframe
df = pd.DataFrame(columns = longitude)

# convert the rows into float data-type
latitude = []

for i in buffer[9:]:
    i = i[:-1]
    i = i.split("\t")

    if i[0] != "":
        if "S" in i[0]:     # if the first entry in the row is not null/blank
            latitude.append(-float(i[0][:-1]))  # append it to latitude list; append -ve for for "S"
            df.loc[-float(i[0][:-1])] = i[1:]   # add the row to the data frame; append -ve for "S" and drop the symbol
        else:
            latitude.append(float(i[0][:-1]))
            df.loc[-float(i[0][:-1])] = i[1:]

print(df.head(5))

temp_col = []
temp_row = []
temp_list = []

temp_col = longitude[0 : ((len(longitude) + 1)) : 10]

for iter1 in temp_col:
    for iter2 in latitude:
        print(df.loc[iter2])

我还提供了数据集here的链接

(下载以.txt结尾的文件,并在与.txt文件相同的目录下运行代码)

我是 numpy、pandas 和 python 的新手,编写这一小段代码对我来说是一项艰巨的任务。如果我能在这方面得到一些帮助,那就太好了。

【问题讨论】:

    标签: python python-3.x pandas numpy


    【解决方案1】:

    欢迎来到 NumPy/Pandas 的世界 :) 它最酷的地方之一是它可以将矩阵上的动作抽象为简单的命令,在绝大多数情况下无需编写循环。

    如果有更多pandorable 代码,您的大量辛勤工作将变得不必要。以下是我尝试重现您所说的内容。我可能误解了,但希望它能让你更接近/为你指明正确的方向。随时要求澄清!

    import pandas as pd
    
    df = pd.read_csv('Aug-2016-potential-temperature-180x188.txt', skiprows=range(7))
    df.columns=['longitude'] #renaming
    df = df.longitude.str.split('\t', expand=True)
    smaller = df.iloc[::10,:] # taking every 10th row
    df.head()
    

    【讨论】:

    • 非常感谢您的回答,尤其是您的链接。非常有用,特别是因为我需要它来完成我的任务!我有两个问题!如何使用您提供的格式迭代数据框?我是一名 Java 程序员,二维数组的 Java 索引在这里不起作用!所以我很困惑。其次,如果我想访问数组的特定元素,可以说数据集 中的行 [1] 和列 [3] 即 21.15628 的元素(在运行代码并执行 df .head(10)) 我该怎么做? df[0][2] 的 Java 方法在这里不起作用!
    • 1) 这个想法是您可以一次完成所有事情,而无需迭代!值得阅读一些关于 NumPy 的内容 - 这样做比一次更改一个数据帧的一个单元格要有效得多。但如果你必须这样做,我认为df.iterrows 可能会有所帮助 2) df.iloc[1,3] !! (一个很好的参考是10 minutes to Pandas)
    【解决方案2】:

    所以如果我理解你的话(只是为了确定): 你有一个巨大的数据集,以纬度和经度作为行和列。 您想从中提取一个子样本来处理它(计算、探索等)。因此,您创建了一个子行列表,并且您希望基于这些行创建一个新的数据框。这是正确的吗?

    如果是这样:

    df['temp_col'] = [ 1 if x%10 == 0 else 0 for x in range(len(longitude))]
    new_df = df[df['temp_col']>0].drop(['temp_col'],axis = 1]
    

    如果您还想删除一些列:

    keep_columns = df.columns.values[0 :len(df.columns) : 10]
    to_be_droped = list(set(df.columns.values) - set(keep_columns))
    new_df = new_df.drop(to_be_droped, axis = 1)
    

    【讨论】:

      猜你喜欢
      • 2021-06-02
      • 2020-06-29
      • 1970-01-01
      • 2021-06-08
      • 2021-09-29
      • 2020-09-21
      • 2014-11-22
      • 1970-01-01
      • 2022-09-24
      相关资源
      最近更新 更多