【问题标题】:removing remaining html tags删除剩余的 html 标签
【发布时间】:2012-03-04 00:25:40
【问题描述】:

我有一个实验室任务,我一直在关注如何删除 html 标记。下面是去除html标签的方法:

public String getFilteredPageContents() {
    String str = getUnfilteredPageContents();
    String temp = "";
    boolean b = false;
    for(int i = 0; i<str.length(); i++) {
        if(str.charAt(i) == '&' || str.charAt(i) == '<') {
            b = true;
        }
        if(b == false) {
            temp += str.charAt(i);
        }
        if(str.charAt(i) == '>' || str.charAt(i) == ';') {
            b = false;
        }
    }
    return temp;
}

这是我的文字最早的形式:

<!DOCTYPE HTML PUBLIC "-//IETF//DTD HTML//EN">
<html>

<head>
<meta http-equiv="Content-Type"
content="text/html; charset=iso-8859-1">
<meta name="GENERATOR" content="Microsoft FrontPage 2.0">
<title>A Shropshire Lad</title>
</head>

<body bgcolor="#008000" text="#FFFFFF" topmargin="10"
leftmargin="20">

<p align="center"><font size="6"><strong></strong></font>&nbsp;</p>
<div align="center"><center>

<pre><font size="7"><strong>A Shropshire Lad
</strong></font><strong>
by A.E. Housman
Published by Dover 1990</strong></pre>
</center></div>

<p><strong>This collection of sixty three poems appeared in 1896.
Many of them make references to Shrewsbury and Shropshire,
however, Housman was not a native of the county. The Shropshire
of his book is a mindscape in which he blends old ballad meters,
classical reminiscences and intense emotional experiences
&quot;recollected in tranquility.&quot; Although they are not
particularly to my taste, their style, simplicity and
timelessness are obvious even to me. Below are two short poems
which amused me, I hope you find them interesting too.</strong></p>

<hr size="8" width="80%" color="#FFFFFF">
<div align="left">

<pre><font size="5"><strong><u>
XIII</u></strong></font><font size="4"><strong>

When I was one-and-twenty
I heard a wise man say,
'Give crowns and pounds and guineas
But not your heart away;</strong></font></pre>
</div><div align="left">

<pre><font size="4"><strong>Give pearls away and rubies
But keep your fancy free.
But I was one-and-twenty,
No use to talk to me.</strong></font></pre>
</div><div align="left">

<pre><font size="4"><strong>When I was one-and-twenty
I heard him say again,
'The heart out of the bosom
Was never given in vain;
'Tis paid with sighs a plenty
And sold for endless rue'
And I am two-and-twenty,
And oh, 'tis true 'tis true.

</strong></font><strong></strong></pre>
</div>

<hr size="8" width="80%" color="#FFFFFF">

<pre><font size="5"><strong><u>LVI . The Day of Battle</u></strong></font><font
size="4"><strong>

'Far I hear the bugle blow
To call me where I would not go,
And the guns begin the song,
&quot;Soldier, fly or stay for long.&quot;</strong></font></pre>

<pre><font size="4"><strong>'Comrade, if to turn and fly
Made a soldier never die,
Fly I would, for who would not?
'Tis sure no pleasure to be shot.</strong></font></pre>

<pre><font size="4"><strong>'But since the man that runs away
Lives to die another day,
And cowards' funerals, when they come,
Are not wept so well at home,</strong></font></pre>

<pre><font size="4"><strong>'Therefore, though the best is bad,
Stand and do the best, my lad;
Stand and fight and see your slain,
And take the bullet in your brain.'</strong></font></pre>

<hr size="8" width="80%" color="#FFFFFF">
</body>
</html>

当在此文本上实现我的方法时:

 charset=iso-8859-1">

A Shropshire Lad







A Shropshire Lad

by A.E. Housman
Published by Dover 1990


This collection of sixty three poems appeared in 1896.
Many of them make references to Shrewsbury and Shropshire,
however, Housman was not a native of the county. The Shropshire
of his book is a mindscape in which he blends old ballad meters,
classical reminiscences and intense emotional experiences
recollected in tranquility. Although they are not
particularly to my taste, their style, simplicity and
timelessness are obvious even to me. Below are two short poems
which amused me, I hope you find them interesting too.
.
.
.

我的问题是:我怎样才能摆脱文本charset=iso-8859-1"&gt; 开头的那个小代码。我无法摆脱那一堆代码?谢谢...

【问题讨论】:

  • 您可以从避免使用首页开始。像这样的工具可以方便地换取正确的代码
  • 避免使用 FrontPage 可能是个好主意。但我认为任务是处理 HTML 代码,不管它来自哪里?

标签: java html tags


【解决方案1】:

我可以看出您的意图是删除类似于 &lt;xxx&gt; 和 &amp;xxx; 的内容。您正在使用变量 b 来记住您当前是否正在跳过内容。

您是否注意到您的算法会跳过&lt;xxx; 和&amp;xxx&gt; 形式的内容?也就是说,&amp; 或 &lt; 将导致跳过开始,&gt; 或 ; 将导致跳过结束,但您不必将 &lt; 与 &gt; 或 &amp; 与;。那么如何实现代码来记住哪个字符开始跳过?

不过,更复杂的是 &amp;xxx; 可以嵌入到 &lt;xxx&gt; 中,如下所示:&lt;p title="&amp;amp;"&gt;

顺便说一句,当字符串很长时,temp += str.charAt(i); 会使您的程序非常慢。看看改用StringBuilder。


这里有一些代码应该可以解决您的问题或几乎可以解决:

import java.util.Stack;

public String getFilteredPageContents() {
    String str = getUnfilteredPageContents();
    StringBuilder() temp = new StringBuilder();

    // The closing character for each thing that we're inside
    Stack<Character> expectedClosing = new Stack<Character>();

    for(int i = 0; i<str.length(); i++) {
        char c = str.charAt(i);
        if(c == '<')
            expectedClosing.push('>');
        else if(c == '&')
            expectedClosing.push(';');

        // Is the current character going to close something?
        else if(!expectedClosing.empty() && c == expectedClosing.peek())
            expectedClosing.pop();

        else {
            // Only add to output if not currently inside something
            if(expectedClosing.empty())
                temp.append(c);
        }
    }
    return temp.toString();
}

【讨论】:

    【解决方案2】:

    这是一项学校作业,但您是否有机会使用格式良好的 HTML 解析器(例如 this)来完成这项工作?

    【讨论】:

      【解决方案3】:

      解决这种情况最优雅的方法可能是使用regular expressions。使用它们,您可以专门搜索标签结构并将它们从输出中删除。

      但是,由于您已经编写了一个程序并且除了您提到的问题之外它工作正常,因此一个快速而肮脏的解决方案可能就足够了。

      我能想到的一件事是应用类似过滤器的算法,逐行扫描文本输出,如果存在则将其删除。就像阅读每一行并检查最后一个字符是否为&gt;。如果是删除该行/用空字符串替换它。在普通文本中不应该有任何&gt; 和句尾,所以你不应该有太多麻烦。

      【讨论】:

        猜你喜欢
        • 2021-08-30
        • 2017-11-12
        • 2017-02-03
        • 1970-01-01
        • 1970-01-01
        • 2021-03-29
        • 2020-03-11
        • 1970-01-01
        • 2012-05-29
        相关资源
        最近更新 更多