【发布时间】:2019-02-08 12:10:37
【问题描述】:
我编写了一个从给定数据集读取数据并将整个 txt 文件转换为 pandas 数据帧的代码(经过一些预处理)
- 纬度代表行并显示在列表中。
- 经度表示列,并显示在单独的列表中。
现在,我想从我创建的原始数据框创建一个更小的数据框(以便更容易理解和解释数据)并执行计算。为此,我通过跳过每 10 个元素创建了一个大小为 18 的较小列。这工作得很好。让我们将此新列称为 new_column。
现在,我想要对每一行进行迭代,并针对第 k 行和新列 j 的每个值,将其添加到新矩阵或数据框。
例如。如果第 10 行和 new_column 12 的值为“x”,我想将这个“x”添加到相同的位置,但在新的数据框(或矩阵)中。
我已经编写了以下代码,但我不知道如何执行让我执行上述操作的部分。
import matplotlib.pyplot as plt
import pandas as pd
import numpy as np
from scipy import interpolate
# open the file for reading
dataset = open("Aug-2016-potential-temperature-180x188.txt", "r+")
# read the file linewise
buffer = dataset.readlines()
# pre-process the data to get the columns
column = buffer[8]
column = column[3 : -1]
# get the longitudes as features
features = column.split("\t")
# convert the features to float data-type
longitude = []
for i in features:
if "W" in features:
longitude.append(-float(i[:-1])) # append -ve sign if "W", drop the "W" symbol
else:
longitude.append(float(i[:-1])) # append +ve sign if "E", drop the "E" symbol
# append the longitude as columns to the dataframe
df = pd.DataFrame(columns = longitude)
# convert the rows into float data-type
latitude = []
for i in buffer[9:]:
i = i[:-1]
i = i.split("\t")
if i[0] != "":
if "S" in i[0]: # if the first entry in the row is not null/blank
latitude.append(-float(i[0][:-1])) # append it to latitude list; append -ve for for "S"
df.loc[-float(i[0][:-1])] = i[1:] # add the row to the data frame; append -ve for "S" and drop the symbol
else:
latitude.append(float(i[0][:-1]))
df.loc[-float(i[0][:-1])] = i[1:]
print(df.head(5))
temp_col = []
temp_row = []
temp_list = []
temp_col = longitude[0 : ((len(longitude) + 1)) : 10]
for iter1 in temp_col:
for iter2 in latitude:
print(df.loc[iter2])
我还提供了数据集here的链接
(下载以.txt结尾的文件,并在与.txt文件相同的目录下运行代码)
我是 numpy、pandas 和 python 的新手,编写这一小段代码对我来说是一项艰巨的任务。如果我能在这方面得到一些帮助,那就太好了。
【问题讨论】:
标签: python python-3.x pandas numpy