【发布时间】:2019-09-18 00:39:13
【问题描述】:
与我的question 类似,我想在 R 中使用 Regex 提取字符串中的字符序列。我想从文本文档中提取部分,从而生成一个数据框,其中每个子部分都被视为自己的矢量,用于进一步的文本挖掘。这是我的示例数据:
chapter_one <- c("One morning, when Gregor Samsa woke from troubled dreams, he found himself transformed in his bed into a horrible vermin.
1 Introduction
He lay on his armour-like back, and if he lifted his head a little he could see his brown belly, slightly domed and divided by arches into stiff sections.
1.1 Futher
The bedding was hardly able to cover it and seemed ready to slide off any moment.
1.1.1 This Should be Part of One Point One
His many legs, pitifully thin compared with the size of the rest of him, waved about helplessly as he looked.
1.2 Futher Fuhter
'What's happened to me?' he thought. It wasn't a dream. His room, a proper human room although a little too small, lay peacefully between its four familiar walls.")
这是我的预期输出:
chapter_id <- (c("1 Introduction", "1.1 Futher", "1.2 Futher Futher"))
text <- (c("He lay on his armour-like back, and if he lifted his head a little he could see his brown belly, slightly domed and divided by arches into stiff sections.", "The bedding was hardly able to cover it and seemed ready to slide off any moment. His many legs, pitifully thin compared with the size of the rest of him, waved about helplessly as he looked.", "'What's happened to me?' he thought. It wasn't a dream. His room, a proper human room although a little too small, lay peacefully between its four familiar walls."))
chapter_one_df <- data.frame(chapter_id, text)
到目前为止我尝试的是这样的:
library(stringr)
regex_chapter_heading <- regex("
[:digit:] # Digit number
# MISSING: Optional dot and optional second digit number
\\s # Space
([[:alpha:]]) # Alphabetic characters (MISSING: can also contain punctuation, as in 'Introduction - A short introduction')
", comments = TRUE)
read.table(text=gsub(regex_chapter_heading,"\\1:",chapter_one),sep=":")
到目前为止,这并没有产生预期的输出 - 因为如前所述,仍然缺少部分正则表达式。非常感谢任何帮助!
【问题讨论】:
-
任何以数字开头的行都是“章节”?也许:
split(txt, grepl("^[1:9]", txt)) -
这是个好主意,但我只想包含直到第二级(1.1 等)的章节。在我的原始文本集中,我有像“2.31.1.1”这样的子章节,我想成为“更大”子章节(在本例中为“2.31”)的一部分。