【发布时间】:2016-12-25 14:41:13
【问题描述】:
我正在使用 GLM(在 Spark 2.0 中使用 ML)对具有一个分类自变量的数据运行模型。我使用StringIndexer 和OneHotEncoder 将该列转换为虚拟变量,然后使用VectorAssembler 将其与连续自变量组合成一列稀疏向量。
如果我的列名是 continuous 和 categorical,其中第一个是浮点列,第二个是表示(在本例中为 8 个)不同类别的字符串:
string_indexer = StringIndexer(inputCol='categorical',
outputCol='categorical_index')
encoder = OneHotEncoder(inputCol ='categorical_index',
outputCol='categorical_vector')
assembler = VectorAssembler(inputCols=['continuous', 'categorical_vector'],
outputCol='indep_vars')
pipeline = Pipeline(stages=string_indexer+encoder+assembler)
model = pipeline.fit(df)
df = model.transform(df)
到目前为止一切正常,我运行模型:
glm = GeneralizedLinearRegression(family='gaussian',
link='identity',
labelCol='dep_var',
featuresCol='indep_vars')
model = glm.fit(df)
model.params
哪些输出:
DenseVector([8440.0573, 3729.449, 4388.9042, 2879.1802, 4613.7646, 5163.3233, 5186.6189, 5513.1392])
这很好,因为我可以验证这些系数基本上是正确的(通过其他来源)。但是,我还没有找到一种将这些系数链接到原始列名的好方法,我需要这样做(我已经为 SO 简化了这个模型;涉及的内容更多。)
StringIndexer 和OneHotEncoder 破坏了列名和系数之间的关系。我找到了一种相当慢的方法:
df[['categorical', 'categorical_index']].distinct()
这给了我一个将字符串名称与数字名称相关联的小数据框,我认为我可以将其与稀疏向量中的键相关联?但是,当您考虑数据的规模时,这非常笨拙且缓慢。
有没有更好的方法来做到这一点?
【问题讨论】:
标签: python pyspark apache-spark-ml