【问题标题】:Strip certain HTML from string从字符串中去除某些 HTML
【发布时间】:2022-01-16 01:09:53
【问题描述】:

我正在使用 ngx-quill,输入正文返回一些 HTML 元素。

例子

<p><strong><em><u>"Soft fingers began to tap the sill of the car window, and the hard fingers tightened on the restless drawing sticks. In the doorways of the sun-beaten tenant houses, women sighed and then shifted feet so that the one that had been down was now on top, and the toes working. Dogs came sniffing near the owner cars and wetted on all four tires one after another. And chickens lay in the sunny dust and fluffed their feathers </u></em></strong></p><p><br></p><p><br></p><p><br></p><p><br></p><p><br></p><p><br></p><p><br></p><p><br></p><p><br></p><p><br></p><p><br></p><p><br></p><p><br></p><p><strong><em>to get the cleansing dust

我想删除所有 HTML 标记,换行段落除外。

当一个帖子有多行/中断时,ngx-quill 会添加几个链接的&lt;p&gt;&lt;/p&gt;&lt;p&gt;&lt;/p&gt;(见上文)

我尝试使用replace 函数来剥离元素,但某些元素(如&lt;u&gt;)并未被删除。另外,如何将具有多个换行符的部分合并为一个换行符

我试过了

post = '<p><strong><em><u>"Soft fingers began to tap the sill of the car window, and the hard fingers tightened on the restless drawing sticks. In the doorways of the sun-beaten tenant houses, women sighed and then shifted feet so that the one that had been down was now on top, and the toes working. Dogs came sniffing near the owner cars and wetted on all four tires one after another. And chickens lay in the sunny dust and fluffed their feathers </u></em></strong></p><p><br></p><p><br></p><p><br></p><p><br></p><p><br></p><p><br></p><p><br></p><p><br></p><p><br></p><p><br></p><p><br></p><p><br></p><p><br></p><p><strong><em>to get the cleansing dust down to the skin. In the little sties the pigs grunted inquiringly over the muddy remnants of the slops.""Soft fingers began to tap the sill of the car window, and the hard fingers tightened on the restless drawing sticks. In the doorways of the sun-beaten tenant houses, women sighed and then shifted feet so that the one that had been down was now on top, and the toes working. Dogs came sniffing near the owner cars and wetted on all four tires one after another. And chickens lay in the sunny dust and fluffed their feathers to get the cleansing dust down to the skin. In the little sties the pigs grunted inquiringly over the muddy remnants of the slops."</em></strong></p>'

function stripElements(post: any) {
    let newPost = post;
    newPost = newPost.replace('<u>', '<span>');
    newPost = newPost.replace('</u>', '</span>');
    newPost = post.replace('<strong>','');
    newPost = newPost.replace('</strong>', '');
    newPost = newPost.replace('<em>', '');
    newPost = newPost.replace('</em>', '');

    newPost = newPost.replace('<p><br></p>', '<p></p>')
    
    return newPost;
}

【问题讨论】:

  • 我不建议使用替换来清理 HTML 字符串。
  • 这能回答你的问题吗? Simple HTML sanitizer in Javascript
  • 不要替换“字符串”?把它变成一个普通的 DOM 树(如果你不能只创建一个临时的 div 之类的,使用documentFragment),然后只使用你需要的 textContent 。然后使用innerHTML 或outerHTML 重新序列化它(如果你真的需要)。
  • 首先,我认为使用.replaceAll()方法会更好
  • 您的代码可以使用 1) 您在第三个 replace => newPost = post.replace('&lt;strong&gt;',''); 处有错字。 -- post 应该是 newPost 和其他人一样。 2) 使用replaceAll 而不是replace 将确保替换所有出现的事件。

标签: javascript string replace


【解决方案1】:

您可以使用DOMParser API 来解析和操作 HTML 代码:

post = '<p><strong><em><u>"Soft fingers began to tap the sill of the car window, and the hard fingers tightened on the restless drawing sticks. In the doorways of the sun-beaten tenant houses, women sighed and then shifted feet so that the one that had been down was now on top, and the toes working. Dogs came sniffing near the owner cars and wetted on all four tires one after another. And chickens lay in the sunny dust and fluffed their feathers </u></em></strong></p><p><br></p><p><br></p><p><br></p><p><br></p><p><br></p><p><br></p><p><br></p><p><br></p><p><br></p><p><br></p><p><br></p><p><br></p><p><br></p><p><strong><em>to get the cleansing dust down to the skin. In the little sties the pigs grunted inquiringly over the muddy remnants of the slops.""Soft fingers began to tap the sill of the car window, and the hard fingers tightened on the restless drawing sticks. In the doorways of the sun-beaten tenant houses, women sighed and then shifted feet so that the one that had been down was now on top, and the toes working. Dogs came sniffing near the owner cars and wetted on all four tires one after another. And chickens lay in the sunny dust and fluffed their feathers to get the cleansing dust down to the skin. In the little sties the pigs grunted inquiringly over the muddy remnants of the slops."</em></strong></p>'

function stripElements(post) {
  const doc = new DOMParser().parseFromString(post, 'text/html');
  doc.querySelectorAll('body :not(p)').forEach(el => el.replaceWith(el.textContent))
  return doc.body.innerHTML;
}

console.log(stripElements(post))

【讨论】:

  • 但是&lt;u&gt; 应该变成&lt;span&gt;... ;)
  • @LouysPatriceBessette 我认为 op 试图完全删除 u,但它没有用,所以他们用 span 进行了故障排除
  • 好的...可能。
【解决方案2】:

规则 #1:不要使用正则表达式操作 HTML。请改用 DOM 解析器。

规则 #2:您可能不想对 DOM 解析器的开销大惊小怪,只想完成工作,并且可能会忽略规则 #1。

因此,如果您愿意,这样的事情可能会奏效:

return post.replace(/<\/?[a-z]+>/gi, m => m.toLowerCase() === '<br>' ? '<p></p>' : '');

我不确定这就是您想要处理换行符的方式,但作为开始,您应该能够根据需要对其进行调整。

【讨论】:

    猜你喜欢
    • 2012-08-07
    • 1970-01-01
    • 2014-11-21
    • 1970-01-01
    • 2014-11-16
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多