【问题标题】:python web scraping csv filepython网页抓取csv文件
【发布时间】:2021-03-05 04:23:22
【问题描述】:

这是我的网页抓取代码,用于获取内容并导出到 csv 文件。我可以知道为什么csv文件的每一行都有间距吗?能解决吗?谢谢!

Python 代码

import requests
from bs4 import BeautifulSoup
import csv

session = requests.session()

payload = {"i0023":"XXXXXX", 
          "i0025":"XXXXXX"
         }
         
session.post("http://192.168.XXX.XXX/checkLogin.cgi",data = payload)

s = session.get("http://192.168.XXX.XXX/m_departmentid.html")

soup = BeautifulSoup(s.text, "html.parser")

table = soup.find('div', attrs={ "class" : "ItemListComponent"})
tbody = table.find_all('tbody')

rows = []

for row in table.find_all('tr'):
    rows.append([val.text for val in row.find_all('td')[0:6]])

with open('test.csv', 'w') as f:
    writer = csv.writer(f)
    writer.writerows(row for row in rows if row)

源代码

<div class="ItemListComponent">
<table>
<thead>
<tr><th rowspan="3" scope="col">Department ID</th><th colspan="5" scope="col">Page Total/Page Restriction</th><th rowspan="3" scope="col"></th></tr>
<tr><th colspan="3" scope="col">Total Prints</th><th colspan="1" scope="col">Color</th><th colspan="1" scope="col">Black & White</th></tr>
<tr><th colspan="1" scope="col">Total</th><th colspan="1" scope="col">Color</th><th colspan="1" scope="col">Black & White</th><th colspan="1" scope="col">Print</th><th colspan="1" scope="col">Print</th></tr>

</thead>
<tbody>
<tr><td>7654321</td><td>11</td><td>0</td><td>11</td><td>0</td><td>11</td><td></td></tr>
<tr><td><a href="/m_departmentid_edit.html?id=100">0000100</a></td><td>0</td><td>0</td><td>0</td><td>0</td><td>0</td><td><input class="ButtonEnable" type="button" value="Delete" title="Delete" onclick="departmentIdDelete(100)"/><input class="ButtonEnable" type="button" value="Clear Count" onclick="departmentIdClear(100)" />
</td></tr>
<tr><td><a href="/m_departmentid_edit.html?id=101">0000101</a></td><td>0</td><td>0</td><td>0</td><td>0</td><td>0</td><td><input class="ButtonEnable" type="button" value="Delete" title="Delete" onclick="departmentIdDelete(101)"/><input class="ButtonEnable" type="button" value="Clear Count" onclick="departmentIdClear(101)" />
</td></tr>
<tr><td><a href="/m_departmentid_edit.html?id=102">0000102</a></td><td>18</td><td>5</td><td>13</td><td>5</td><td>13</td><td><input class="ButtonEnable" type="button" value="Delete" title="Delete" onclick="departmentIdDelete(102)"/><input class="ButtonEnable" type="button" value="Clear Count" onclick="departmentIdClear(102)" />
</td></tr>
<tr><td><a href="/m_departmentid_edit.html?id=103">0000103</a></td><td>0</td><td>0</td><td>0</td><td>0</td><td>0</td><td><input class="ButtonEnable" type="button" value="Delete" title="Delete" onclick="departmentIdDelete(103)"/><input class="ButtonEnable" type="button" value="Clear Count" onclick="departmentIdClear(103)" />
</td></tr>
<tr><td><a href="/m_departmentid_edit.html?id=104">0000104</a></td><td>0</td><td>0</td><td>0</td><td>0</td><td>0</td><td><input class="ButtonEnable" type="button" value="Delete" title="Delete" onclick="departmentIdDelete(104)"/><input class="ButtonEnable" type="button" value="Clear Count" onclick="departmentIdClear(104)" />
</td></tr>

1

【问题讨论】:

    标签: python beautifulsoup python-requests export-to-csv


    【解决方案1】:

    您将其打开为“wb”,即写入字节。改为“w”打开它。

    【讨论】:

    • 经过测试,现在可以导出为csv文件了。但是,当我打开 csv 文件时,数据不正确。只要标题正确。谢谢。
    • 我会进一步提供帮助,但它没有可以从中提取的示例位置,因此美丽的汤不会返回任何东西。但如果我不得不猜测,那是因为你的行文本每次通过循环时都会被重置,所以你错过了除了最后一个循环之外的所有内容。
    • 感谢您的帮助。期待帮助我解决它非常感谢。
    【解决方案2】:

    您需要对字符串进行编码以将其转换为字节对象。

    for row in soup.select(".ItemListComponent tbody tr")[1:215]:
        row_text = [x.text.encode() for x in row.find_all("td")]
        print(",".join(row_text))
    

    【讨论】:

    • 经过测试,还是不行。像这样的错误消息:AttributeError:'str'对象没有属性'decode'。谢谢。
    • 糟糕,我看错了。您需要编码以将字符串转换为字节。我已经更新了代码。
    • 对不起。这个错误是什么意思:TypeError: sequence item 0: expected str instance, bytes found?谢谢。
    • 和join函数有关系吗?请指教。谢谢。
    • 您在哪个行号收到此错误?能否请您发送错误回溯>
    【解决方案3】:

    谢谢大家。最后,我找到了解决在 csv writer 中添加换行参数时缺少的问题的解决方案。

    代码

    session = requests.session()
    
    payload = {"i0023":"XXXXX", 
              "i0025":"XXXXX"
             }
             
    session.post("http://192.168.XXX.XXX/checkLogin.cgi",data = payload)
    
    s = session.get("http://192.168.XXX.XXX/m_departmentid.html")
    
    soup = BeautifulSoup(s.text, "html.parser")
    
    table = soup.find('div', attrs={ "class" : "ItemListComponent"})
    table_tbody = table.find('tbody')
    
    rows = []
     
    for row in table.find_all('tr'):
        rows.append([val.text for val in row.find_all('td')])   
    
    
    with open(("\test.csv"), 'w', newline='') as f:
        writer = csv.writer(f)
        writer.writerows(row for row in rows if row)
    

    【讨论】:

      猜你喜欢
      • 2014-06-20
      • 1970-01-01
      • 2019-03-12
      • 1970-01-01
      • 2021-01-28
      • 2013-11-14
      • 1970-01-01
      • 2022-09-27
      • 2019-11-12
      相关资源
      最近更新 更多