【问题标题】:Parsing specific content using BeautifulSoup使用 BeautifulSoup 解析特定内容
【发布时间】:2016-02-12 23:01:31
【问题描述】:

我希望从网站中提取梦幻足球信息。 我可以编写足够的代码来获得以下输出,但我真正想要的是以下信息:

"fullName":"Justin Forsett"
"pointsSEASON":75

谁能帮助解释如何隔离这些项目并将它们写入例如 csv 文件?

[<div class="mod-content" id="fantasy-content">{"averagePoints":9.4,"percentOwned":98.6,"pointsSEASON":75,"seasonOutlook":{"outlook":"Forsett finished 2014 as fantasy's No. 8 RB, so why aren't we higher on him? Well, it's difficult to reconcile what we know about his size (5-8, 197), age (30 in October) and career with the 1,529 scrimmage yards he racked up as Baltimore's surprise starter. Forsett had never even eclipsed 1,000 total yards in any of his six previous seasons. Yet his quickness and vision were consistently excellent last year, and new OC Marc Trestman loves throwing to RBs. Lorenzo Taliaferro and rookie Javorius Allen loom as heftier options, and some kind of rotation could develop. But Forsett will get the benefit of the doubt in Week 1.","seasonId":2015,"date":"Wed May 20"},"positionRank":18,"playerId":11467,"percentChange":-0.2,"averageDraftPosition":42.5,"fullName":"Justin Forsett","mostRecentNews":{"news":null,"spin":"The Jaguars have allowed the second-fewest yards per carry (3.4) in the league, but have ceded one rushing score per game in the process. Forsett will need a good deal of volume to overcome a quietly tough matchup, but we're trusting the workload will be enough.","date":"Tue Nov 10"},"totalPoints":75,"projectedPoints":13.957546548,"projectedDifference":4.582546548}</div>]

【问题讨论】:

    标签: python xml-parsing beautifulsoup


    【解决方案1】:

    您要查找的标签文本似乎是 JSON 格式。您已成功获取div 标签,但现在您必须提取 JSON,然后提取您想要的信息。这是您需要添加到代码中的内容。

    import json
    
    rawJSONString = {originaltag}.get_text()
    JSONString = json.loads(rawJSONString)
    print(JSONString['fullName'])
    print(JSONString['pointsSEASON'])
    

    {originaltag} 是您在上面打印的标签,因为您没有显示您的代码,所以我无法运行它。相反,我运行了以下代码

    string = '{"averagePoints":9.4,"percentOwned":98.6,"pointsSEASON":75,"seasonOutlook":{"outlook":"Forsett finished 2014 as fantasys No. 8 RB, so why arent we higher on him? Well, its difficult to reconcile what we know about his size (5-8, 197), age (30 in October) and career with the 1,529 scrimmage yards he racked up as Baltimores surprise starter. Forsett had never even eclipsed 1,000 total yards in any of his six previous seasons. Yet his quickness and vision were consistently excellent last year, and new OC Marc Trestman loves throwing to RBs. Lorenzo Taliaferro and rookie Javorius Allen loom as heftier options, and some kind of rotation could develop. But Forsett will get the benefit of the doubt in Week 1.","seasonId":2015,"date":"Wed May 20"},"positionRank":18,"playerId":11467,"percentChange":-0.2,"averageDraftPosition":42.5,"fullName":"Justin Forsett","mostRecentNews":{"news":null,"spin":"The Jaguars have allowed the second-fewest yards per carry (3.4) in the league, but have ceded one rushing score per game in the process. Forsett will need a good deal of volume to overcome a quietly tough matchup, but were trusting the workload will be enough.","date":"Tue Nov 10"},"totalPoints":75,"projectedPoints":13.957546548,"projectedDifference":4.582546548}'
    s = json.loads(string)
    print(s['fullName'])
    print(s['pointsSEASON'])
    

    得到这个输出

    Justin Forsett
    75
    

    编辑添加:Here 是有关写入 csv 文件的信息。

    【讨论】:

    • 感谢您的回复,我尝试将您提供的内容添加到我的代码中,但它显示了与“get_text()”部分相关的错误消息。在我的代码中,上面打印的标签的变量是“标签”。我所做的只是用标签替换 {originaltag} 并产生以下错误:ResultSet 对象没有属性'get_text()'。对我所缺少的任何帮助都会很棒。再次感谢。
    • 请将代码添加到您的原始问题中,这样更容易调试。当你有一个字符串而不是一个标签时,通常会发生该错误,因此请尝试在不使用 .get_text() 方法的情况下运行它。
    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 2017-10-29
    • 1970-01-01
    • 2019-05-24
    • 2017-10-28
    • 2019-10-29
    • 2011-04-29
    • 2012-02-05
    相关资源
    最近更新 更多