Python爬虫爬取一个网页上的图片地址实例代码

442 阅读 0 评论 292 点赞

我是靠谱客的博主矮小奇迹，这篇文章主要介绍Python爬虫爬取一个网页上的图片地址实例代码，现在分享给大家，希望可以做个参考。

本文实例主要是实现爬取一个网页上的图片地址，具体如下。

读取一个网页的源代码：

import urllib.request
def getHtml(url):
  html=urllib.request.urlopen(url).read()
  return html
print(getHtml(http://image.baidu.com/search/flip?tn=baiduimage&ie=utf-8&word=%E5%A3%81%E7%BA%B8&ct=201326592&lm=-1&v=flip))

利用正则表达式爬取一个网页上的图片地址：

import re
import urllib.request
def getHtml(url):
  html=urllib.request.urlopen(url).read()
  return html
def getImg(html):
  r=r'"thumbURL":"(http://img.+?\.jpg)"' #定义正则
  imglist=re.findall(r,html)
  return imglist
html=str(getHtml("http://image.baidu.com/search/flip?tn=baiduimage&ie=utf-8&word=%E5%A3%81%E7%BA%B8&ct=201326592&lm=-1&v=flip"))
print(getImg(html))

运行结果：