回答您的问题“我想我的底线问题是如何访问网页上的第二个 <li> 的所有内容?”使用相对现代、得到良好支持并内置于 PHP 中的 API:
<?php
$url = "https://waset.org/conferences-in-january-2022-in-tokyo";
libxml_use_internal_errors(true);
$dom = new DomDocument();
$dom->loadHtmlFile($url);
$lists = $dom->getElementsByTagName("ul");
$items = $lists[1]->getElementsByTagName("li");
foreach ($items as $item) {
// clean up extra whitespace
$text = preg_replace("/\s+/", " ", trim($item->textContent));
echo "$text\n------\n";
}
输出:
ICA 2022: Aeroponics Conference, Tokyo (Jan 07-08, 2022)
------
ICAA 2022: Agroforestry and Applications Conference, Tokyo (Jan 07-08, 2022)
------
ICAAAA 2022: Applied Aerodynamics, Aeronautics and Astronautics Conference, Tokyo (Jan 07-08, 2022)
------
ICAAAE 2022: Aquatic Animals and Aquaculture Engineering Conference, Tokyo (Jan 07-08, 2022)
------
ICAAC 2022: Advances in Astronomical Computing Conference, Tokyo (Jan 07-08, 2022)
------
...
另外值得注意的是,会议名称在<a> 元素中,地点在其中的<span> 中,日期紧随其后。使用它,您可以相当简单地提取数据:
function getNodeText(\DomNode $node): string
{
$return = "";
foreach($node->childNodes as $child) {
if ($child->nodeName === "#text") {
$return .= trim($child->nodeValue);
}
}
return $return;
}
foreach ($items as $item) {
$conference = getNodeText($item->getElementsByTagName("a")[0]);
$location = getNodeText($item->getElementsByTagName("span")[0]);
$date = getNodeText($item);
echo "------\n$conference | $location | $date\n";
}
输出:
------
ICA 2022: Aeroponics Conference, | Tokyo | (Jan 07-08, 2022)
------
ICAA 2022: Agroforestry and Applications Conference, | Tokyo | (Jan 07-08, 2022)
------
ICAAAA 2022: Applied Aerodynamics, Aeronautics and Astronautics Conference, | Tokyo | (Jan 07-08, 2022)
------
ICAAAE 2022: Aquatic Animals and Aquaculture Engineering Conference, | Tokyo | (Jan 07-08, 2022)
------
ICAAC 2022: Advances in Astronomical Computing Conference, | Tokyo | (Jan 07-08, 2022)
...