【发布时间】:2018-11-28 23:02:02
【问题描述】:
我正在尝试使用列名绘制某些基于树的模型的特征重要性。我正在使用 Pyspark。
因为我也有文本分类变量和数字变量,所以我不得不使用类似这样的管道方法 -
- 使用字符串索引器来索引字符串列
- 对所有列使用一个热编码器
-
使用向量组装器创建包含特征向量的特征列
来自docs 的一些示例代码,用于步骤 1、2、3 -
from pyspark.ml import Pipeline from pyspark.ml.feature import OneHotEncoderEstimator, StringIndexer, VectorAssembler categoricalColumns = ["workclass", "education", "marital_status", "occupation", "relationship", "race", "sex", "native_country"] stages = [] # stages in our Pipeline for categoricalCol in categoricalColumns: # Category Indexing with StringIndexer stringIndexer = StringIndexer(inputCol=categoricalCol, outputCol=categoricalCol + "Index") # Use OneHotEncoder to convert categorical variables into binary SparseVectors # encoder = OneHotEncoderEstimator(inputCol=categoricalCol + "Index", outputCol=categoricalCol + "classVec") encoder = OneHotEncoderEstimator(inputCols= [stringIndexer.getOutputCol()], outputCols=[categoricalCol + "classVec"]) # Add stages. These are not run here, but will run all at once later on. stages += [stringIndexer, encoder] numericCols = ["age", "fnlwgt", "education_num", "capital_gain", "capital_loss", "hours_per_week"] assemblerInputs = [c + "classVec" for c in categoricalColumns] + numericCols assembler = VectorAssembler(inputCols=assemblerInputs, outputCol="features") stages += [assembler] # Create a Pipeline. pipeline = Pipeline(stages=stages) # Run the feature transformations. # - fit() computes feature statistics as needed. # - transform() actually transforms the features. pipelineModel = pipeline.fit(dataset) dataset = pipelineModel.transform(dataset) -
最终训练模型
在训练和评估之后,我可以使用“model.featureImportances”来获得特征排名,但是我没有获得特征/列名,而只是特征编号,像这样 -
print dtModel_1.featureImportances (38895,[38708,38714,38719,38720,38737,38870,38894],[0.0742343395738,0.169404823667,0.100485791055,0.0105823115814,0.0134236162982,0.194124862158,0.437744255667])
如何将其映射回初始列名和值?这样我就可以绘图了?**
【问题讨论】:
标签: apache-spark pyspark apache-spark-sql apache-spark-mllib