溫馨提示×

您好,登錄后才能下訂單哦!

密碼登錄×
登錄注冊(cè)×
其他方式登錄
點(diǎn)擊 登錄注冊(cè) 即表示同意《億速云用戶服務(wù)條款》

通過selenium實(shí)現(xiàn)的京東商品爬取

發(fā)布時(shí)間:2020-07-05 12:48:26 來源:網(wǎng)絡(luò) 閱讀:13066 作者:AESCR 欄目:編程語(yǔ)言

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as ec
from lxml import etree
import csv
import requests,re,time
#搜索的商品名稱
shopname="Python設(shè)計(jì)模式"
#聲明瀏覽器對(duì)象
browser=webdriver.Chrome()
browser.get("https://www.jd.com")
#查找節(jié)點(diǎn)
inputtext = browser.find_element_by_class_name('text')
#輸入數(shù)據(jù)
inputtext.send_keys(shopname)
#提交
btn = browser.find_element_by_class_name('button')
btn.click()
#搜索后的頁(yè)面
#顯式等待
wait = WebDriverWait(browser, 10)
wait.until(ec.title_contains(shopname))
with open(shopname+".csv",'a') as f:
    wr= csv.DictWriter(f,['name','price','shop'])
    wr.writeheader()
    while True:
        #判斷是否為反爬蟲機(jī)制窗體 是否正常
        if len(browser.window_handles)>1:
            handles=browser.window_handles[1]
            browser.switch_to_window(handles)
            browser.close()
        # 滾動(dòng)條
        browser.execute_script("window.scrollTo(0, document.body.scrollHeight)")
        wait.until(ec.presence_of_element_located((By.CLASS_NAME, 'pn-next')))
        # 爬取內(nèi)容
        html = etree.HTML(browser.page_source)
        # 讀取每個(gè)商品
        shops = html.xpath('//div[contains(@class,"gl-i-wrap")]')
        # 下一頁(yè)
        npage =html.xpath('//a[@class="pn-next disabled"]/em//text()')
        for shop in shops:
            name = shop.xpath('.//div[contains(@class,"p-name")]//em//text()')
            name = "".join(name)
            price = shop.xpath('.//div[contains(@class,"p-price")]//i//text()')
            price = "".join(price)
            sname = shop.xpath('.//div[contains(@class,"p-shop")]//a//@title')
            sname = "".join(sname)
            if sname.strip() == '':
                sname = "京東自營(yíng)"
            wr.writerow({'name':name,'price':price,'shop':sname})

        if len(npage)>0:
            break
        try:
            pbtn = browser.find_element_by_class_name("pn-next")
            pbtn.click()
        except:
            pass

    browser.close()
向AI問一下細(xì)節(jié)

免責(zé)聲明:本站發(fā)布的內(nèi)容(圖片、視頻和文字)以原創(chuàng)、轉(zhuǎn)載和分享為主,文章觀點(diǎn)不代表本網(wǎng)站立場(chǎng),如果涉及侵權(quán)請(qǐng)聯(lián)系站長(zhǎng)郵箱:is@yisu.com進(jìn)行舉報(bào),并提供相關(guān)證據(jù),一經(jīng)查實(shí),將立刻刪除涉嫌侵權(quán)內(nèi)容。

AI