<strong id="m7twe"></strong><ul id="m7twe"><th id="m7twe"></th></ul>

<tfoot id="m7twe"><optgroup id="m7twe"></optgroup></tfoot>
<u id="m7twe"><tr id="m7twe"></tr></u>

<menuitem id="m7twe"></menuitem>

<strong id="m7twe"><delect id="m7twe"><menu id="m7twe"></menu></delect></strong>

<strike id="m7twe"></strike>

<strike id="m7twe"><table id="m7twe"></table></strike>

溫馨提示×

溫馨提示×

您好，登錄后才能下訂單哦！

密碼登錄×

忘記密碼？

登錄注冊×

獲取短信驗(yàn)證碼

其他方式登錄

點(diǎn)擊登錄注冊即表示同意《億速云用戶服務(wù)條款》

用戶登錄×

賬戶密碼登錄

請使用微信掃描上方二維碼

使用幫助

請求超時(shí)！

請點(diǎn)擊重新獲取二維碼

怎么用Python來理清楚紅樓夢里的關(guān)系

發(fā)布時(shí)間：2021-10-28 18:05:16 來源：億速云閱讀：144 作者：柒染欄目：編程語言

本篇文章為大家展示了怎么用Python來理清楚紅樓夢里的關(guān)系，內(nèi)容簡明扼要并且容易理解，絕對能使你眼前一亮，通過這篇文章的詳細(xì)介紹希望你能有所收獲。

最近把紅樓夢又抽空看了一遍，古典中的經(jīng)典，我真無法用言辭贊美她。今天，想跟大家一起用 Python 來理一理紅樓夢中的的那些關(guān)系

不要問我為啥是紅樓夢，而不是水滸三國或西游，都是經(jīng)典，但我個(gè)人還是更喜歡偏古典的書，紅樓夢也是我多次反復(fù)品讀的為數(shù)不多的小說，對它的感情也是最深的。

好了好了這些都不重要，重要的是我們今天要用Python來理紅樓夢的關(guān)系！

數(shù)據(jù)準(zhǔn)備

紅樓夢 TXT 文件一份
金陵十二釵 + 賈寶玉人物名稱列表
人物列表內(nèi)容如下：

寶玉 nr

黛玉 nr

寶釵 nr

湘云 nr

鳳姐 nr

李紈 nr

元春 nr

迎春 nr

探春 nr

惜春 nr

妙玉 nr

巧姐 nr

秦氏 nr

這份列表，同時(shí)也是為了做分詞時(shí)使用，后面的 nr 就是人名的意思。

數(shù)據(jù)處理

讀取數(shù)據(jù)并加載詞典

 with open("紅樓夢.txt", encoding='gb18030') as f:
 honglou = f.readlines()
 jieba.load_userdict("renwu_forcut")
 renwu_data = pd.read_csv("renwu_forcut", header=-1)
 mylist = [k[0].split(" ")[0] for k in renwu_data.values.tolist()]

這樣，我們就把紅樓夢讀取到了 honglou 這個(gè)變量當(dāng)中，同時(shí)也通過 load_userdict 將我們自定義的詞典加載到了 jieba 庫中。

對文本進(jìn)行分詞處理并提取

tmpNames = []
 names = {}
 relationships = {}
 for h in honglou:
 h.replace("賈妃", "元春")
 h.replace("李宮裁", "李紈")
 poss = pseg.cut(h)
 tmpNames.append([])
 for w in poss:
 if w.flag != 'nr' or len(w.word) != 2 or w.word not in mylist:
 continue
 tmpNames[-1].append(w.word)
 if names.get(w.word) is None:
 names[w.word] = 0
 relationships[w.word] = {}
 names[w.word] += 1

首先，因?yàn)槲闹?quot;賈妃", “元春”，“李宮裁”, “李紈” 混用嚴(yán)重，所以這里直接做替換處理。

然后使用 jieba 庫提供的 pseg 工具來做分詞處理，會返回每個(gè)分詞的詞性。

之后做判斷，只有符合要求且在我們提供的字典列表里的分詞，才會保留。

一個(gè)人每出現(xiàn)一次，就會增加一，方便后面畫關(guān)系圖時(shí)，人物 node 大小的確定。

對于存在于我們自定義詞典的人名，保存到一個(gè)臨時(shí)變量當(dāng)中 tmpNames。

處理人物關(guān)系

 for name in tmpNames:
 for name1 in name:
 for name2 in name:
 if name1 == name2:
 continue
 if relationships[name1].get(name2) is None:
 relationships[name1][name2] = 1
 else:
 relationships[name1][name2] += 1

對于出現(xiàn)在同一個(gè)段落中的人物，我們認(rèn)為他們是關(guān)系緊密的，每同時(shí)出現(xiàn)一次，關(guān)系增加1.

保存到文件

 with open("relationship.csv", "w", encoding='utf-8') as f:
 f.write("Source,Target,Weight\n")
 for name, edges in relationships.items():
 for v, w in edges.items():
 f.write(name + "," + v + "," + str(w) + "\n")
 with open("NameNode.csv", "w", encoding='utf-8') as f:
 f.write("ID,Label,Weight\n")
 for name, times in names.items():
 f.write(name + "," + name + "," + str(times) + "\n")

文件1：人物關(guān)系表，包含首先出現(xiàn)的人物、之后出現(xiàn)的人物和一同出現(xiàn)次數(shù)
文件2：人物比重表，包含該人物總體出現(xiàn)次數(shù)，出現(xiàn)次數(shù)越多，認(rèn)為所占比重越大。

制作關(guān)系圖表

使用 pyecharts 作圖

def deal_graph():
 relationship_data = pd.read_csv('relationship.csv')
 namenode_data = pd.read_csv('NameNode.csv')
 relationship_data_list = relationship_data.values.tolist()
 namenode_data_list = namenode_data.values.tolist()
 nodes = []
 for node in namenode_data_list:
 if node[0] == "寶玉":
 node[2] = node[2]/3
 nodes.append({"name": node[0], "symbolSize": node[2]/30})
 links = []
 for link in relationship_data_list:
 links.append({"source": link[0], "target": link[1], "value": link[2]})
 g = (
 Graph()
 .add("", nodes, links, repulsion=8000)
 .set_global_opts(title_opts=opts.TitleOpts(title="紅樓人物關(guān)系"))
 )
 return g

首先把兩個(gè)文件讀取成列表形式
對于“寶玉”，由于其占比過大，如果統(tǒng)一進(jìn)行縮放，會導(dǎo)致其他人物的 node 過小，展示不美觀，所以這里先做了一次縮放

最后得出的關(guān)系圖

怎么用Python來理清楚紅樓夢里的關(guān)系

上述內(nèi)容就是怎么用Python來理清楚紅樓夢里的關(guān)系，你們學(xué)到知識或技能了嗎？如果還想學(xué)到更多技能或者豐富自己的知識儲備，歡迎關(guān)注億速云行業(yè)資訊頻道。

向AI問一下細(xì)節(jié)

推薦閱讀：

免責(zé)聲明：本站發(fā)布的內(nèi)容（圖片、視頻和文字）以原創(chuàng)、轉(zhuǎn)載和分享為主，文章觀點(diǎn)不代表本網(wǎng)站立場，如果涉及侵權(quán)請聯(lián)系站長郵箱：is@yisu.com進(jìn)行舉報(bào)，并提供相關(guān)證據(jù)，一經(jīng)查實(shí)，將立刻刪除涉嫌侵權(quán)內(nèi)容。

上一篇新聞：
一行Python命令搞定前期數(shù)據(jù)探索性的方法是什么
下一篇新聞：
Mysql數(shù)據(jù)分組排名實(shí)現(xiàn)的示例分析

猜你喜歡

AI
助
手

產(chǎn)品服務(wù)

地區(qū)劃分

專題活動

幫助支持

關(guān)于我們

售后咨詢

7*24小時(shí)在線電話：400-100-2938

7*24小時(shí)在線 QQ：800811969

關(guān)注億速云

億速云公眾號

手機(jī)網(wǎng)站二維碼

<strike id="6u22u"></strike>

<dfn id="6u22u"><progress id="6u22u"><th id="6u22u"></th></progress></dfn>

<menuitem id="6u22u"></menuitem>

<div id="6u22u"><em id="6u22u"><del id="6u22u"></del></em></div>

<delect id="6u22u"></delect>