python爬取盘搜的有效链接实现代码-侯体宗的博客

python爬取盘搜的有效链接实现代码
Python / 管理员发布于 8年前 373

因为盘搜搜索出来的链接有很多已经失效了，影响找数据的效率，因此想到了用爬虫来过滤出有效的链接，顺便练练手~

这是本次爬取的目标网址http://www.pansou.com，首先先搜索个python，之后打开开发者工具，

可以发现这个链接下的json数据就是我们要爬取的数据了，把多余的参数去掉，

剩下的链接格式为http://106.15.195.249:8011/search_new?q=python&p=1，q为搜索内容，p为页码

以下是代码实现：

import requestsimport jsonfrom multiprocessing.dummy import Pool as ThreadPoolfrom multiprocessing import Queueimport sysheaders = {  "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/75.0.3770.100 Safari/537.36"}q1 = Queue()q2 = Queue()urls = [] # 存取url列表# 读取urldef get_urls(query):  # 遍历50页  for i in range(1,51):    # 要爬取的url列表，返回值是json数据，q参数是搜索内容，p参数是页码    url = "http://106.15.195.249:8011/search_new?&q=%s&p=%d" % (query,i)    urls.append(url)# 获取数据def get_data(url):  print("开始加载，请等待...")  # 获取json数据并把json数据转换为字典  resp = requests.get(url, headers=headers).content.decode("utf-8")  resp = json.loads(resp)  # 如果搜素数据为空就抛出异常停止程序  if resp['list']['data'] == []:    raise Exception  # 遍历每一页数据的长度  for num in range(len(resp['list']['data'])):    # 获取百度云链接    link = resp['list']['data'][num]['link']    # 获取标题    title = resp['list']['data'][num]['title']    # 访问百度云链接，判断如果页面源代码中有“失效时间：”这段话的话就表明链接有效，链接无效的页面是没有这段话的    link_content = requests.get(link, headers=headers).content.decode("utf-8")    if "失效时间：" in link_content:      # 把标题放进队列1      q1.put(title)      # 把链接放进队列2      q2.put(link)      # 写入csv文件      with open("wangpanziyuan.csv", "a+", encoding="utf-8") as file:        file.write(q1.get()+","+q2.get() + "\n")  print("ok")if __name__ == '__main__':  # 括号内填写搜索内容  get_urls("python")  # 创建线程池  pool = ThreadPool(3)  try:    results = pool.map(get_data, urls)  except Exception as e:    print(e)  pool.close()  pool.join()  print("退出")

总结

以上所述是小编给大家介绍的python爬取盘搜的有效链接实现代码希望对大家有所帮助，如果大家有任何疑问请给我留言，小编会及时回复大家的。在此也非常感谢大家对站的支持！
如果你觉得本文对你有帮助，欢迎转载，烦请注明出处，谢谢！

上一条：
python 字符串追加实例
下一条：
python将字符串list写入excel和txt的实例

0条评论 (评论内容有缓存机制,请悉知!)

最新最热

近期评论
test1 在
opencode + Oh-my-openagent,我的第一个免费的ai编程智能体管家:Sisyphus中评论 test..
122 在
学历：一种延缓就业设计，生活需求下的权衡之选中评论工作几年后，报名考研了，到现在还没认真学习备考，迷茫中。作为一名北漂互联网打工人..
Zita 在
Google AI Studio升级全栈 vibe coding体验，可直接构建带登录和数据库的应用中评论 111222..
123 在
Clash for Windows作者删库跑路了，github已404中评论按理说只要你在国内，所有的流量进出都在监控范围内，不管你怎么隐藏也没用，想搞你分..
原梓番博客在
在Laravel框架中使用模型Model分表最简单的方法中评论好久好久都没看友情链接申请了，今天刚看，已经添加。..

Top