python3+selenium获取页面加载的所有静态资源文件链接操作-侯体宗的博客

python3+selenium获取页面加载的所有静态资源文件链接操作
Python / 管理员发布于 8年前 622

软件版本：

python 3.7.2

selenium 3.141.0

pycharm 2018.3.5

具体实现流程如下，废话不多说，直接上代码：

from selenium import webdriverfrom selenium.webdriver.chrome.options import Optionsfrom selenium.webdriver.common.desired_capabilities import DesiredCapabilitiesd = DesiredCapabilities.CHROMEchrome_options = Options()#使用无头浏览器chrome_options.add_argument('--headless')chrome_options.add_argument('--user-agent=Mozilla/5.0 (Windows NT 10.0; WOW64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/71.0.3578.98 Safari/537.36')#浏览器启动默认最大化chrome_options.add_argument("--start-maximized");#该处替换自己的chrome驱动地址browser = webdriver.Chrome("D://googleDever//chromedriver.exe",chrome_options=chrome_options,desired_capabilities=d)browser.set_page_load_timeout(150)browser.get("https://www.xxx.com")#静态资源链接存储集合urls = []#获取静态资源有效链接for log in browser.get_log('performance'): if 'message' not in log:continue log_entry = json.loads(log['message']) try:#该处过滤了data:开头的base64编码引用和document页面链接if "data:" not in log_entry['message']['params']['request']['url'] and 'Document' not in log_entry['message']['params']['type']:urls.append(log_entry['message']['params']['request']['url']) except Exception as e:pass print(urls)

打印结果为页面渲染时加载的静态资源文件链接：

[http://www.xxx.com/aaa.js,http://www.xxx.com/css.css]

以上代码为selenium获取页面加载过程中预加载的各类静态资源文件链接，使用该功能获取到链接后，使用其他插件进行可对资源进行下载！

补充知识：在idea 中python import sys，import requests 报错

File->Project Structure
project -> sdk -> new -> ok

设置编译参数（主要是设置和检查Python JDK是否正确）

以上这篇python3+selenium获取页面加载的所有静态资源文件链接操作就是小编分享给大家的全部内容了，希望能给大家一个参考，也希望大家多多支持。

上一条：
Python插件机制实现详解
下一条：
python3 sleep 延时秒毫秒实例

0条评论 (评论内容有缓存机制,请悉知!)

最新最热

近期评论
test1 在
opencode + Oh-my-openagent,我的第一个免费的ai编程智能体管家:Sisyphus中评论 test..
122 在
学历：一种延缓就业设计，生活需求下的权衡之选中评论工作几年后，报名考研了，到现在还没认真学习备考，迷茫中。作为一名北漂互联网打工人..
Zita 在
Google AI Studio升级全栈 vibe coding体验，可直接构建带登录和数据库的应用中评论 111222..
123 在
Clash for Windows作者删库跑路了，github已404中评论按理说只要你在国内，所有的流量进出都在监控范围内，不管你怎么隐藏也没用，想搞你分..
原梓番博客在
在Laravel框架中使用模型Model分表最简单的方法中评论好久好久都没看友情链接申请了，今天刚看，已经添加。..

Top