【发布时间】:2017-02-22 16:08:14
【问题描述】:
有人建议我使用 BeautifulSoup 从 HTML 中删除具有特定 id 的标签。例如,删除<div id=needDelete>...</div> 下面是我的代码,但似乎不能正常工作:
import os, re
from bs4 import BeautifulSoup
cwd = os.getcwd()
print ('Now you are at this directory: \n' + cwd)
# find files that have an extension with HTML
Files = os.listdir(cwd)
print Files
def func(file):
for file in os.listdir(cwd):
if file.endswith('.html'):
print ('HTML files are \n' + file)
f = open(file, "r+")
soup = BeautifulSoup(f, 'html.parser')
matches = str(soup.find_all("div", id="jp-post-flair"))
#The soup.find_all part should be correct as I tested it to
#print the matches and the result matches the texts I want to delete.
f.write(f.read().replace(matches,''))
#maybe the above line isn't correct
f.close()
func(file)
你能帮忙检查一下哪个部分的代码有误吗?也许我应该如何处理它? 非常感谢!!
【问题讨论】:
标签: python html tags beautifulsoup