【问题标题】:Define words in Python在 Python 中定义单词
【发布时间】:2015-07-29 07:38:40
【问题描述】:

这似乎与此重复:Python define a word?

但是,这并不是因为我试图在我的代码中实现该答案(适用于该线程的 OP,但不适用于我)。

这是我的功能:

def define_word(user_define_input):
    srch = str(user_define_input[1])
    output_word=urllib.request.urlopen("http://dictionary.reference.com/browse/"+srch+"?s=t")
    output_word=output_word.read()
    items=re.findall('<meta name="description" content="'+".*$",output_word,re.MULTILINE)
    for output_word in items:
        y=output_word.replace('<meta name="description" content="','')
        z=y.replace(' See more."/>','')
        m=re.findall('at Dictionary.com, a free online dictionary with pronunciation, synonyms and translation. Look it up now! "/>',z)
        if m==[]:
            if z.startswith("Get your reference question answered by Ask.com"):
                print ("Word not found!")
            else:
                print (z)
    else:
        print ("Word not found!")

注意:

>>> print (user_define_input) #to show what is in the list
>>> define <word entered> #prints out the list, in this case, the program ignores user_define_input[0] and looks for [1] which is the targeted word

此外,这包含一些 HTML :/ 抱歉,但这就是其他答案使用的内容。

所以,当我尝试使用这个时的错误:

File "/Users/******/GitHub/Multitool/functions.py", line 104, in define_word
items=re.findall('<meta name="description" content="'+".*$",output_word,re.MULTILINE)
File "/Library/Frameworks/Python.framework/Versions/3.4/lib/python3.4/re.py", line 210, in findall
return _compile(pattern, flags).findall(string)
TypeError: can't use a string pattern on a bytes-like object

注意: functions.py 的第 104 行是:

items=re.findall('<meta name="description" content="'+".*$",output_word,re.MULTILINE)

re.py的第210行是这个函数的最后一行:

def findall(pattern, string, flags=0):
    """Return a list of all non-overlapping matches in the string.

    If one or more capturing groups are present in the pattern, return a list of groups; this will be a list of tuples if the pattern
has more than one group.

Empty matches are included in the result."""
    return _compile(pattern, flags).findall(string) #line 210

如果有什么不清楚的地方,请告诉我(我不确定要为这个添加什么标签:/)。并提前感谢您:) 随意更改任何内容,甚至重写整个内容,但请确保使用变量/列表:

  • define_word(用于函数名称)
  • user_define_input

如果您想查看 git,请转到此链接:https://github.com/DarkLeviathanz/Multitool.git

添加:

output_word = output_word.decode()

或改变

output_word = output_word.read().decode('iso-8859-2')

在输入时给出了这个:定义测试:

Test definition, the means by which the presence, quality, or genuineness of anything is determined; a means of trial.<meta property="og:url" content="http://dictionary.reference.com/browse/test"/><link rel="shortcut icon" href="http://static.sfdict.com/dictcloud/favicon.ico"/><!--[if lt IE 9]><link rel="respond-proxy" id="respond-proxy" href="http://static.sfdict.com/app/respondProxy-d7e5f.html" /><![endif]--><!--[if lt IE 9]><link rel="respond-redirect" id="respond-redirect" href="http://dictionary.reference.com/img/respond.proxy.gif" /><![endif]--><link rel="search" type="application/opensearchdescription+xml" href="http://dictionary.reference.com/opensearch_desc.xml" title="Dictionary.com"/><link rel="publisher" href="https://plus.google.com/117428481782081853923"/><link rel="canonical" href="http://dictionary.reference.com/browse/test"/><link rel="stylesheet" href="http://dictionary.reference.com/drc/css/bootstrap.min-93899.css" type="text/css" media="all"/><link rel="stylesheet" href="http://dictionary.reference.com/drc/css/combinedSerp-8c61a.css" type="text/css" media="all"/><script type="text/javascript">var searchURL="http://dictionary.reference.com/browse/%40%40queryText%40%40?s=t";var CTSParams={"infix":"","clkpage":"dic","clksite":"dict","clkld":0};</script>
Word not found!

【问题讨论】:

    标签: python html function python-3.x dictionary


    【解决方案1】:
    output_word = output_word.decode()
    

    将字节转换为字符串。


    更新

    这是聊天中脚本的最后一个状态(仍然远非完美......):

    import requests
    from lxml import html
    
    def define_word(word):
        response = requests.get(
            "http://dictionary.reference.com/browse/{}?s=t".format(word))
        tree = html.fromstring(response.text)
        title = tree.xpath('//title/text()')
        print(title)
        defs = tree.xpath('//div[@class="def-content"]/text()')
        # print(defs)
    
        defs = ''.join(defs)
        defs = defs.split('\n')
        defs = [d for d in defs if d]
        for d in defs:
            print(d)
    
    define_word('python')
    

    【讨论】:

    • kk,但是在定义打印后我得到了很多html文本。
    • 是的,现在您需要从 html 中提取您需要的数据。我建议您使用像docs.python.org/3/library/html.parser.html 或crummy.com/software/BeautifulSoup/bs4/doc 之类的html 解析器...
    • 我只是建议您使用一种用于从 html 中提取数据的工具 - 而不是正则表达式。正则表达式也可以工作,但我敢打赌它比用于此目的的工具更痛苦。
    • 终于!它至少可以工作:)谢谢,顺便说一句,在我的情况下用user_define_input[1]替换单词:)。
    【解决方案2】:

    urllib.request.urlopen().read() 返回一个字节串。异常表示当将 Python 字符串应用于字节字符串时,不能将其用作正则表达式模式。

    字节字符串(通常)是编码的 unicode 字符串,在这种情况下,它看起来像 UTF-8 编码的数据。因此,您需要将字节字符串解码为 Python 字符串,以便将其用作正则表达式模式:

    output_word = urllib.request.urlopen("http://dictionary.reference.com/browse/"+srch+"?s=t")
    output_word = output_word.read().decode('utf8')
    

    这应该可以为您解决问题。

    您确实需要知道要使用什么编码。这可以通过查看Content-Type 响应标头来完成,对于此URL,它是Content-Type: text/html; charset=UTF-8。或者,由于这是 HTML 内容,您可以查找 &lt;meta http-equiv="Content-type" ... 标记。

    最后,您可以使用 requests 库来为您处理这个问题:

    import requests
    r = requests.get("http://dictionary.reference.com/browse/"+srch+"?s=t")
    output_word = r.text
    

    【讨论】:

    • 它解决了这个问题,但是在打印定义后显示了很多奇怪的 html 内容。
    • 对不起,那可能是因为我本来说数据编码为ISO-8859-2(由chardet.detect()报道),但现在看来确实是UTF-8。跨度>
    • and... ImportError: No module named 'requests'?这很奇怪......但它的 python 3.4.2 如此 idk,也许这改变了。而且,额外的 html 代码(有问题添加)可能不是因为编码。可能是因为 html 的 find 部分存在一些问题。
    • requests 是第 3 方库。您可以使用pip install requests 安装它或查看此处:docs.python-requests.org/en/latest/user/install/#install
    • 是的,我明白了。我还使用该命令安装了 lxml/hmtl :) ty
    【解决方案3】:

    经过一些更改,这是我坚持使用的代码,尽管它仍然存在一些缺陷。

    def define_word(user_define_input):
        try:
            response = requests.get("http://dictionary.reference.com/browse/{}?s=t".format(user_define_input[1]))
        except IndexError:
            print("You have not entered a word!")
            return
        tree = html.fromstring(response.text)
        title = tree.xpath('//title/text()')
        print(title)
        print("\n")
        defs = tree.xpath('//div[@class="def-content"]/text()')
        defs = ''.join(defs)
        defs = defs.replace("() ", "")
        defs = defs.split('\n')
        defs = [d for d in defs if d]
        for d in defs:
            print(d)
    

    这将用户输入拆分为一个包含两个项目的列表:

    def split_line_test(user_input):
        global user_define_input
        user_define_input = user_input.split()
        if (user_define_input[0] == "define"): #define is user_define_input[0] while user_define_input[1] is the word that will be searched up
            return True
        if (user_define_input[0] == "weather"): #you can ignore this, it is for my other function
            return True
        return False
    

    非常感谢帮助我修复代码的人 :)

    【讨论】:

      猜你喜欢
      • 1970-01-01
      • 2017-04-02
      • 2011-01-30
      • 1970-01-01
      • 1970-01-01
      • 2021-05-10
      • 1970-01-01
      • 2013-10-21
      • 2023-04-02
      相关资源
      最近更新 更多