【问题标题】:processing xml file with hadoop using python使用 python 使用 hadoop 处理 xml 文件
【发布时间】:2012-10-27 09:40:25
【问题描述】:

我正在使用 python 和 hadoop 来处理一个 xml 文件,我有以下格式的 xml 文件

temporary.xml

<report>
<report-name name="ALL_TIME_KEYWORDS_PERFORMANCE_REPORT"/>
<date-range date="All Time"/>
  <table>
    <columns>
       <column name="campaignID" display="Campaign ID"/>
       <column name="adGroupID" display="Ad group ID"/>
       <column name="keywordID" display="Keyword ID"/>
       <column name="keyword" display="Keyword"/>
    </columns>
    <row campaignID="79057390" adGroupID="3451305670" keywordID="3000000" keyword="Content"/>
    <row campaignID="79057390" adGroupID="3451305670" keywordID="3000000" keyword="Content"/>
    <row campaignID="79057390" adGroupID="3451305670" keywordID="3000000" keyword="Content"/>
    <row campaignID="79057390" adGroupID="3451305670" keywordID="3000000" keyword="Content"/>
  </table>
</report>

现在我要做的就是处理上面的 xml 文件,然后将数据保存到 MSSQL 数据库中。

ma​​pper.py代码

import sys
import cStringIO
import xml.etree.ElementTree as xml

if __name__ == '__main__':
    buff = None
    intext = False
    for line in sys.stdin:
        line = line.strip()
        if line.find("<row>") != -1:
            intext = True
            buff = cStringIO.StringIO()
            buff.write(line)
        elif line.find("</>") != -1:
            intext = False
            buff.write(line)
            val = buff.getvalue()
            buff.close()
            buff = None
            print val

在这里,我要做的就是从row tags 获取数据,即campaignID,adgroupID,keywordID,keyword 的值并打印它们,作为reducer.py 的输入(其中包含将数据保存在数据库中的代码)。

我查看了一些示例,但标签类似于&lt;tag&gt; &lt;/tag&gt;,但就我而言,我只有&lt;row/&gt;

但是我上面的代码不起作用/没有打印任何东西,任何人都可以更正我的代码并添加必要的 python 代码以从行标签中获取值/数据(我对 hadoop 非常非常陌生),所以下次会扩展代码。

【问题讨论】:

    标签: python xml hadoop


    【解决方案1】:

    你考虑过使用 xpath 吗?它是一种迷你语言,可用于绕过 xml 树。它可以在 python 中轻松使用。

    http://docs.python.org/2/library/xml.etree.elementtree.html 可能对你有用

    您可能还想查看Need Help using XPath in ElementTree

    我会这样做(这是有效的 Python 代码。我在 Python3.2 中对其进行了测试。适用于您的示例 xml):

    import xml.etree.ElementTree as xml #you had this line in your code. I am not using any tool you  do not have access to in your script
    
    def get_row_attributes(the_xml_as_a_string):
        """
        this function takes xml as a string. 
        It can work with xml that looks like your included example xml.
        This function returns a list of dictionaries. Each dictionary is made up of the attributes of each row. So the result looks like:
         [
              {attribute_name:value_for_first_row,attribute_name:value_for_first_row...},
              {attribute_name:value_for_second_row,attribute_name:value_for_second_row...},
              etc
         ]
        """
        tree = xml.fromstring(the_xml_as_a_string)
        rows = tree.findall('table/row')  # 'table/row' is xpath. it means get all the rows in all the tables
        return [row.attrib for row in rows]
    

    要使用这个函数,读入标准并建立一个字符串。致电get_row_attributes(the_xml_as_a_string)

    生成的字典包含您请求的信息(行的属性)。

    所以现在我们有

    1. 从标准输入中读取内容
    2. 获取所有行的所有信息

    全部使用完全正常的python

    最后要做的就是将它写入您的其他进程。如果您需要这部分的帮助,请提供有关数据应采用何种格式以及应放在何处的信息

    【讨论】:

    • :感谢您宝贵的回复,但我正在通过 haddop 尝试此操作并希望实际保存到数据库中,所以请您添加代码以从 hadoop 和 python 的行标签中获取值
    • @shivakrishna:为了清晰起见,我添加了一些 cmets。如果您需要其他任何内容,请具体说明
    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 2014-06-07
    • 2012-03-04
    • 2011-11-02
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多