Article Topic Learning Download Q&A Programming Dictionary Game Recent Updates

简体中文(ZH-CN) English(EN) 繁体中文(ZH-TW) 日本語(JA) 한국어(KO) Melayu(MS) Français(FR) Deutsch(DE)

Home > Backend Development > PHP Tutorial > body text

关于preg_match_all的抓取,该如何解决

WBOY

Release： 2016-06-13 12:55:11

Original

884 people have browsed it

关于preg_match_all的抓取

<div><br />
<h1>标题1</h1><br />
<p>内容1</p><br />
<p>内容2</p><br />
<h1>标题2</h1><br />
<p>内容1</p><br />
<p>内容2</p><br />
<p>内容3</p><br />
<p>内容4</p><br />
<h1>标题3</h1><br />
<p>内容1</p><br />
<p>内容2</p><br />
<p>内容3</p><br />
</div>

Copy after login

我要用preg_match_all()来循环获取从

到下一个

之前的内容即

标题1

内容1

内容2

－－－－－－－－－－－－

标题2

内容1

内容2

内容3

内容4

－－－－－－－－－－－－

标题3

内容1

内容2

内容3

我想过用

preg_match_all('/<h1>[\w\W]*<(h1|\/div)/U',$html, $out)

Copy after login

但这样抓，会隔一个就跳过，因为第二个的

已经被第一个用了。

------解决方案--------------------

preg_match_all('/<div>(.*)<\/div>/is', $str, $m);<br />
$m = explode('<h1>', substr($m[1][0], 5));<br />
foreach($m as $x)<br />
    echo htmlspecialchars ("<h1>$x") . '<br/>';

Copy after login

Related labels：

gt lt match nbsp

source：php.cn

Previous article： CURL中文乱码解决方案 Next article：软件工程师2013新年增值计划，转自php100

Statement of this Website

The content of this article is voluntarily contributed by netizens, and the copyright belongs to the original author. This site does not assume corresponding legal responsibility. If you find any content suspected of plagiarism or infringement, please contact admin@php.cn