php只抓取网页头的方法:1、使用get_headers()函数;2、使用http_response_header方法;3、使用stream_get_meta_data()函数;4、使用php curl来获取网页头即可。
本文操作环境:windows7系统、php7.1版、dell g3电脑
php如何只抓取网页头?
php获取网页header信息的4种方法
php获取网页header信息的方法多种多样,就php语言来说,我知道的方法有4种, 下面逐一献上。
方法一:使用get_headers()函数
推荐指数: ★★★★★
get_header方法最简单只要两行代码即可搞定。如下:
$thisurl = "http://www.lao8.org/";print_r(get_headers($thisurl, 1));
得到的结果为:
array( [0] => http/1.1 200 ok [cache-control] => max-age=86400 [content-length] => 76102 [content-type] => text/html [content-location] => http://www.lao8.org/index.html [last-modified] => fri, 19 jul 2013 03:52:30 gmt [accept-ranges] => bytes [etag] => "50bc48643384ce1:5cb3" [server] => microsoft-iis/6.0 [x-powered-by] => asp.net [date] => fri, 19 jul 2013 09:06:39 gmt [connection] => close)
方法二:使用http_response_header
推荐指数: ★★★
http_response_headerf方法也很简单,仅三行:
$thisurl = "http://www.lao8.org";$html = file_get_contents($thisurl ); print_r($http_response_header);
得到的结果为:
array( [0] => http/1.1 200 ok [1] => cache-control: max-age=86400 [2] => content-length: 76102 [3] => content-type: text/html [4] => content-location: http://www.lao8.org/index.html [5] => last-modified: fri, 19 jul 2013 03:52:30 gmt [6] => accept-ranges: bytes [7] => etag: "50bc48643384ce1:5cb3" [8] => server: microsoft-iis/6.0 [9] => x-powered-by: asp.net [10] => date: fri, 19 jul 2013 09:06:41 gmt [11] => connection: close)
方法三:使用stream_get_meta_data()函数
推荐指数: ★★★
使用stream_get_meta_data()代码也只需三行:
$thisurl = "http://www.lao8.org/";$fp = fopen($thisurl, 'r'); print_r(stream_get_meta_data($fp));
得到的结果为:
array( [wrapper_data] => array ( [0] => http/1.1 200 ok [1] => cache-control: max-age=86400 [2] => content-length: 76102 [3] => content-type: text/html [4] => content-location: http://www.lao8.org/index.html [5] => last-modified: fri, 19 jul 2013 03:52:30 gmt [6] => accept-ranges: bytes [7] => etag: "50bc48643384ce1:5cb3" [8] => server: microsoft-iis/6.0 [9] => x-powered-by: asp.net [10] => date: fri, 19 jul 2013 09:06:41 gmt [11] => connection: close ) [wrapper_type] => http [stream_type] => tcp_socket [mode] => r+ [unread_bytes] => 1086 [seekable] => [uri] => http://www.lao8.org/ [timed_out] => [blocked] => 1 [eof] => )
第四种方法: 使用php的高级函数 curl()来获取
推荐指数: ★★★★
上面的三种方法能获取一般的网页header信息,如果想要获取更详细的header信息比如网页是否启用了gzip压缩。这时候可以用php的高级函数curl()来获取。
使用curl获得header可以检测gzip压缩
先贴出代码:
<?php$szurl = 'http://www.lao8.org/';$curl = curl_init();curl_setopt($curl, curlopt_url, $szurl);curl_setopt($curl, curlopt_header, 1); //输出header信息curl_setopt($curl, curlopt_returntransfer, 1); //不显示网页内容curl_setopt($curl, curlopt_encoding, ''); //允许执行gzip$data=curl_exec($curl); if(!curl_errno($curl)){ $info = curl_getinfo($curl); $httpheadersize = $info['header_size']; //header字符串体积 $pheader = substr($data, 0, $httpheadersize); //获得header字符串 $split = array("rn", "n", "r"); //需要格式化header字符串 $pheader = str_replace($split, '<br>', $pheader); //使用<br>换行符格式化输出到网页上 echo $pheader;}?>
输出结果如下:
http/1.1 200 okcache-control: max-age=86400content-length: 15189content-type: text/htmlcontent-encoding: gzipcontent-location: http://www.lao8.org/index.htmllast-modified: fri, 19 jul 2013 03:52:28 gmtaccept-ranges: bytesetag: "0268633384ce1:5cb3"vary: accept-encodingserver: microsoft-iis/6.0x-powered-by: asp.netdate: fri, 19 jul 2013 09:27:21 gmt
可以看到使用curl获取到的header信息多了这行:content-encoding: gzip,网页启用了gzip压缩。
推荐学习:《php视频教程》
以上就是php如何只抓取网页头的详细内容。