【问题标题】:Is there any way to transform a string according to text row above?有没有办法根据上面的文本行转换字符串?
【发布时间】:2020-07-14 23:36:03
【问题描述】:

我有一个很长的数组列表,其中包括属行(以大写字母开头,例如:ACHNATHES)和物种行(以大写字母开头,然后是一个点,例如:A.)我需要根据上面的内容进行一些转换文本。下面你就很容易理解了:

ACHNANTHES
A. brevipes
A. coarctata
A. cocconeiformis
A. gibberula
A. lacunarum 
A. lineariformis
A. longipes
A. nollii
A. parvula
A. petersenii
A. pyrenaicum
A. stolida
A. thermalis
A. trinodis
A. wellsiae
PLATESSA
P. conspicua
P. montana
P. salinarum
ACHNANTHIDIUM
A. affine
A. deflexum
A. exiguum
A. exile
A. lanceolatum
A. minutissimum
A. minutum
A. thermale

我想把它改成这样:

ACHNANTHES
Achantes brevipes
Achantes coarctata
Achantes cocconeiformis
Achantes gibberula
Achantes lacunarum
Achantes lineariformis
Achantes longipes
Achantes nollii
Achantes parvula
Achantes petersenii
Achantes pyrenaicum
Achantes stolida
Achantes thermalis
Achantes trinodis
Achantes wellsiae
PLATESSA
Platessa conspicua
Platessa montana
Platessa salinarum
ACHNANTHIDIUM
Achanthidium affine
Achanthidium deflexum
Achanthidium exiguum
Achanthidium exile
Achanthidium lanceolatum

认为我需要在 PHP 中使用 while 或 foreach,但我不知道该怎么做。请帮忙。

@mickmackmusa 如你所愿,是我数组的一部分:

array (
  0 => 'ACHNANTHES Bory, Dict. Class. Hist. Nat. 1: 79 (1822). / SUCINCIĞI.',
  1 => 'A. brevipes C.Agardh, Syst. Alg.: 1 (1824). / Küçük sucıncığı.',
  2 => 'A. coarctata (Bréb. ex W.Sm.) Grunow, Syn. Diat. Belg.: expl. pl. XXVI: ş. 17 (1880). / Dar sucıncığı.',
  3 => 'A. cocconeiformis Mann, U.S. Nat. Mus., Bull. 6: 182 (1925). / Top sucıncığı.',
  4 => 'A. gibberula Grunow, Kongl. Svenska Vetensk.-Akad. Handl. 17(2): 121 (1880). / Kambur sucıncığı.',
  5 => 'A. lacunarum Hust., Bacillariophyta (Diatomeae) Zweite Auflage, Süsswass.-Fl. Mitteleurop. 10: 205 (1930). / Delikli sucıncığı.',
  6 => 'A. lineariformis Lange-Bert., Biblioth. Diatomol. 27: 7, 134 pl. (1993). / Düz sucıncığı.',
  7 => 'A. longipes C.Agardh, Syst. Alg.: 1 (1824). / Boylu sucıncığı.',
  8 => 'A. nollii Bock, Nachrichtendes Naturwiss. Museums Stadt Aschaffenburg 38: 1 (1953). / Yaban sucıncığı.',
  9 => 'A. parvula Kütz., Bacillarien: 76, pl. 21: ş. 5 (1844). / Saf sucıncığı.',
  10 => 'A. petersenii Hust., Rabenhorst’s Krypt.-Fl. Deutschl.: 179, ş 10-14 (1937). / Bal sucıncığı.',
  11 => 'A. pyrenaicum (Hust.) H.Kobayasi, Nova Hedwigia 65(1-4): 148, ş. 1-18 (1997). / Garip sucıncığı.',
  12 => 'A. stolida (Krasske) Krasske, Ann. Acad. Sc. Fenn., ser. A, Biol. 14: 78 (1949). / Alık sucıncığı.',
  13 => 'A. thermalis (Rabenh.) Schoenfeld, Diat. German.: 122 (1907). / Sıcak sucıncığı.',
  14 => 'A. trinodis (Ralfs) Grunow, Syn. Diatom. Belg.: pl. XXVII: ş. 50 (1880). / Üç sucıncığı.',
  15 => 'A. wellsiae Reimer, Monogr. Acad. Nat. Sci. Philadelphia 1: 16 (1966). / El sucıncığı.',
  16 => 'PLATESSA Lange-Bert., Süsswass.-Fl. Mitteleuropa 2: 443 (2004). / SUTANESİ.',
  17 => 'P. conspicua (Ant.Mayer) Lange-Bert., Süsswass.-Fl. Mitteleuropa 2: 445 (2004). / Küt sutanesi.',
  18 => 'P. montana (Krasske) Lange-Bert., Süsswass.-Fl. Mitteleuropa 2: 445 (2004). / Dağ sutanesi.',
  19 => 'P. salinarum (Grunow) Lange-Bert. / Sutanesi.',
  20 => 'ACHNANTHIDIUM Kütz., Bacillarien: 75 (1844). / SUÇUBUĞU.',
  21 => 'A. affine (Grunow) Czarn., Mem. Calif. Acad. Sci. 17: 156 (1994). / Hoş suçubuğu.',
  22 => 'A. deflexum Kingston, Diatom Res. 15(2): 409 (2000). / Kıvrık suçubuğu.',
  23 => 'A. exiguum (Grunow) Czarnecki, Mem. Calif. Acad. Sci. 17: 155 (1994). / Delikli suçubuğu.',
  24 => 'A. exile (Kütz.) Heiberg, Conspect. Diatom. Dan.: 119 (1863). / Bitik suçubuğu.',
  25 => 'A. lanceolatum (Bréb.) Kütz., Bot. Zeitung 4(14): 247 (1846). / Uzun suçubuğu.',
  26 => 'A. minutissimum (Kütz.) Czarn., Mem. Calif. Acad. Sci. 17: 155 (1994). / Cüce suçubuğu.',
  27 => 'A. minutum Cleve, Fl. Fenn. 8(2): 1 (1891). / Bodur suçubuğu.',
  28 => 'A. thermale Rabenh., Fl. Eur. Alg. 1: 107 (1864). / Sıcak suçubuğu.',
  29 => 'EUCOCCONEIS Cleve ex Meister, Beitr. Kryptogamenfl. Schweiz IV(1): 95 (1912). / SUESNEĞİ.',
  30 => 'E. flexella (Kütz.) Meister, Beitr. Kryptogamenfl. Schweiz IV(1): 95 (1912). / Suesneği.',
  31 => 'E. laevis (Østrup) Lange-Bert., Iconogr. Diatomol. 6: 46 (1999). / Pek suesneği.',
  32 => 'E. quadratarea (Østrup) Lange-Bert., Iconogr. Diatomol. 6: 48 (1999). / Dört suesneği.',

【问题讨论】:

  • 这些字符串是单独的数组项吗?即,你有类似$species = ['ACHNANTHES', 'A. brevipes', 'A. coarctata', ...];
  • 是的,每一行都是一个数组索引,如[0]、[1]...
  • 此列表是否已排序?为什么不将其存储为诸如[ARCHANTHES => [ brevipes,...] 之类的键值,当您要访问它们时,将键降低并为其获取以下值
  • 我不知道该怎么做,这是一个很长的列表,比如 4000 行
  • 作为对您的问题的编辑,请以最真实的数组形式提供您的数据的小型、真实的表示形式。使用var_export() 并将该文本复制粘贴到您的问题中。 @Shraun

标签: php regex for-loop while-loop do-while


【解决方案1】:

我有点失望,我没有足够的细节让你一直到查询过程,所以我只会改变你的元素值。

  1. 建立一个分组字符串 -- Genus 变量。进入循环前设置为null
  2. 在您进行迭代时,通过提取第一个单词来确定当前行是否是一个属值,然后检查它是否仅由大写字母组成。
    • 如果是,则将其缓存为新的分组值并将其存储到输出数组中
    • 如果不是,则将格式化的“属种”字符串推入结果数组中

我喜欢正则表达式,但由于您的数据已经被拆分为元素,因此在此任务中使用正则表达式没有任何好处。

代码:(Demo)

$result = [];
$currentGenus = null;
foreach ($array as $line) {
    $firstWord = strstr($line, ' ', true);
    if (ctype_upper($firstWord)) {
        $currentGenus = $firstWord;
        $result[] = $firstWord;
    } else {
        $result[] = ucfirst(strtolower($currentGenus)) . ' ' . explode(' ', $line, 3)[1];
    }
}
var_export($result);

输出:

array (
  0 => 'ACHNANTHES',
  1 => 'Achnanthes brevipes',
  2 => 'Achnanthes coarctata',
  3 => 'Achnanthes cocconeiformis',
  4 => 'Achnanthes gibberula',
  5 => 'Achnanthes lacunarum',
  6 => 'Achnanthes lineariformis',
  7 => 'Achnanthes longipes',
  8 => 'Achnanthes nollii',
  9 => 'Achnanthes parvula',
  10 => 'Achnanthes petersenii',
  11 => 'Achnanthes pyrenaicum',
  12 => 'Achnanthes stolida',
  13 => 'Achnanthes thermalis',
  14 => 'Achnanthes trinodis',
  15 => 'Achnanthes wellsiae',
  16 => 'PLATESSA',
  17 => 'Platessa conspicua',
  18 => 'Platessa montana',
  19 => 'Platessa salinarum',
  20 => 'ACHNANTHIDIUM',
  21 => 'Achnanthidium affine',
  22 => 'Achnanthidium deflexum',
  23 => 'Achnanthidium exiguum',
  24 => 'Achnanthidium exile',
  25 => 'Achnanthidium lanceolatum',
  26 => 'Achnanthidium minutissimum',
  27 => 'Achnanthidium minutum',
  28 => 'Achnanthidium thermale',
  29 => 'EUCOCCONEIS',
  30 => 'Eucocconeis flexella',
  31 => 'Eucocconeis laevis',
  32 => 'Eucocconeis quadratarea',
)

【讨论】:

    【解决方案2】:

    为了解决您的问题,在您考虑任何 PHP 语言细节之前,我会首先考虑您将使用什么样的逻辑。大多数通用编程语言(例如 PHP)在字符串操作方面几乎可以做任何你需要的事情,所以不要担心你现在将如何实现你的逻辑。

    我认为在这种情况下使用正则表达式库会有点矫枉过正。有很多方法可以解决你的问题,而且通常有比我脑海中第一次出现的更好的方法,但我会回顾一下我脑海中第一次出现的逻辑。

    首先,我将回顾一些重要的假设。暗示属行仅包含字母,而种行将以字母开头,然后是点。我还假设了三个新事物:

    1. 除了属行和种行之外,没有其他类型的行
    2. 属行至少有两个字符长
    3. 第一行是一个属名。

    所有这些假设都应该是正确的,如果是,那么这个解决方案将起作用。这是我的英文逻辑:

    Declare a variable that will be a string that keeps track of your current genus name 
    
    For each line (AKA for each string in your array), do this chunk of code:
      See if the second letter of the current line is not a dot
        If it is not, this line is your current genus name: change
          your current genus name variable to the current line
      BUT... if the second letter of the current line IS a dot
        This is a species line, and we will need to transform it, and to do that...
        Make a new string that is the current line with the first two characters cut off
        Make a new string copy of your current genus name, but where it just 
          starts with a capital instead of being all-caps
        Make a new string, which is those two strings you just made put together
        Replace the current line with that newest string you just made
    

    现在,我不会给你一个彻底的解决方案,因为如果我剥夺了你这个学习机会,Stack Overflow 会恨我,但我会让你知道一些关于如何解决这个问题的有用语法。

    foreach 循环 https://www.w3schools.in/php/looping/foreach/

    字符串 https://www.php.net/language.types.string(搜索“按字符访问和修改字符串”)

    if 和 else 语句 https://www.w3schools.com/php/php_if_else.asp

    子字符串 https://www.php.net/manual/en/function.substr.php

    有用的字符串大小写函数 https://www.javatpoint.com/php-string-strtolower-function

    字符串连接 https://www.php.net/manual/en/language.operators.string.php

    附: - 一个真正好的解决方案将具有错误处理,例如如果属名只有一个字符长,或者只有返回字符的行等,但为了简单起见,我没有在这个解决方案中做到这一点。此答案应该适合您的目的,请记住,错误处理是一种很好的做法,并且会为您省去很多麻烦。

    【讨论】:

      【解决方案3】:

      我很高兴被证明是错误的,但我不认为使用 PCRE 正则表达式引擎可以通过简单的替换获得所需的结果。

      假设字符串是

      ACHNANTHES
      A. brevipes
      A. coarctata
      A. cocconeiformis
      PLATESSA
      P. conspicua
      P. montana
      P. salinarum
      

      如果你把线颠倒来获得

      P. salinarum
      P. montana
      P. conspicua
      PLATESSA
      A. cocconeiformis
      A. coarctata
      A. brevipes
      ACHNANTHES
      

      你可以使用正则表达式

      ^[A-Z]\.(?=\s+[a-z]+\s*(?:[A-Z]\.\s+[a-z]+\s*)*([A-Z]+)\s*$)
      

      获取匹配并将每个替换为捕获组的内容,获取

      PLATESSA salinarum
      PLATESSA montana
      PLATESSA conspicua
      PLATESSA
      ACHNANTHES cocconeiformis
      ACHNANTHES coarctata
      ACHNANTHES brevipes
      ACHNANTHES
      

      此时通过反转这些行可以获得所需的结果:

      ACHNANTHES
      ACHNANTHES brevipes
      ACHNANTHES coarctata
      ACHNANTHES cocconeiformis
      PLATESSA
      PLATESSA conspicua
      PLATESSA montana
      PLATESSA salinarum
      

      Demo

      以下操作由 PHP 的正则表达式引擎 PCRE 执行。

      ^                # match beginning of line
      [A-Z]\.          # match uc ltr then '.' 
      (?=              # begin non-cap grp
        \s+[a-z]+\s*   # match 1+ whtspaces, 1+ lc ltrs, 0+ whtspaces
        (?:            # begin non-cap grp
          [A-Z]\.      # match line begin with uc ltr then '.'
          \s+[a-z]+\s* # match 1+ whtspaces, 1+ lc ltrs, 0+ whtspaces 
        )              # end non-cap grp
        *              # execute non-cap grp 0+ times
        ([A-Z]+)       # match 1+ uc ltrs in cap grp 1
        \s*            # match 0+ whtspaces
        $              # match end of line
      )                # end positive lookahead      
      

      【讨论】:

      • 我相信我可以证明你错了,但我会等到 OP 用更准确的输入表示和所需的确切输出更新问题。
      • @mickmackusa,我希望你这样做(真的)。对于澳大利亚人来说,你的手柄很奇怪。
      • ...一定是阅读障碍
      猜你喜欢
      • 2011-02-23
      • 2022-08-12
      • 2014-06-02
      • 2023-01-30
      • 1970-01-01
      • 2012-03-09
      • 1970-01-01
      • 2019-09-30
      • 1970-01-01
      相关资源
      最近更新 更多