python 全文检索引擎详解-侯体宗的博客

python 全文检索引擎详解
Python / 管理员发布于 8年前 237

python 全文检索引擎详解

最近一直在探索着如何用Python实现像百度那样的关键词检索功能。说起关键词检索，我们会不由自主地联想到正则表达式。正则表达式是所有检索的基础，python中有个re类，是专门用于正则匹配。然而，光光是正则表达式是不能很好实现检索功能的。

python有一个whoosh包，是专门用于全文搜索引擎。

whoosh在国内使用的比较少，而它的性能还没有sphinx/coreseek成熟，不过不同于前者，这是一个纯python库，对python的爱好者更为方便使用。具体的代码如下

安装

输入命令行 pip install whoosh

需要导入的包有:

fromwhoosh.index import create_infromwhoosh.fields import *fromwhoosh.analysis import RegexAnalyzerfromwhoosh.analysis import Tokenizer,Token

中文分词解析器

class ChineseTokenizer(Tokenizer):  """  中文分词解析器  """  def __call__(self, value, positions=False, chars=False,         keeporiginal=True, removestops=True, start_pos=0, start_char=0,         mode='', **kwargs):    assert isinstance(value, text_type), "%r is not unicode "% value    t = Token(positions, chars, removestops=removestops, mode=mode, **kwargs)    list_seg = jieba.cut_for_search(value)    for w in list_seg:      t.original = t.text = w      t.boost = 0.5      if positions:        t.pos = start_pos + value.find(w)      if chars:        t.startchar = start_char + value.find(w)        t.endchar = start_char + value.find(w) + len(w)      yield tdef chinese_analyzer():  return ChineseTokenizer()

构建索引的函数

@staticmethod  def create_index(document_dir):    analyzer = chinese_analyzer()    schema = Schema(titel=TEXT(stored=True, analyzer=analyzer), path=ID(stored=True),content=TEXT(stored=True, analyzer=analyzer))    ix = create_in("./", schema)    writer = ix.writer()    for parents, dirnames, filenames in os.walk(document_dir):      for filename in filenames:        title = filename.replace(".txt", "").decode('utf8')        print title        content = open(document_dir + '/' + filename, 'r').read().decode('utf-8')        path = u"/b"        writer.add_document(titel=title, path=path, content=content)    writer.commit()

检索函数

 @staticmethod  def search(search_str):    title_list = []    print 'here'    ix = open_dir("./")    searcher = ix.searcher()    print search_str,type(search_str)    results = searcher.find("content", search_str)    for hit in results:      print hit['titel']      print hit.score      print hit.highlights("content", top=10)      title_list.append(hit['titel'])    print 'tt',title_list    return title_list

感谢阅读，希望能帮助到大家，谢谢大家对本站的支持！

上一条：
python 网络编程详解及简单实例
下一条：
Python处理PDF及生成多层PDF实例代码

0条评论 (评论内容有缓存机制,请悉知!)

最新最热

近期评论
test1 在
opencode + Oh-my-openagent,我的第一个免费的ai编程智能体管家:Sisyphus中评论 test..
122 在
学历：一种延缓就业设计，生活需求下的权衡之选中评论工作几年后，报名考研了，到现在还没认真学习备考，迷茫中。作为一名北漂互联网打工人..
Zita 在
Google AI Studio升级全栈 vibe coding体验，可直接构建带登录和数据库的应用中评论 111222..
123 在
Clash for Windows作者删库跑路了，github已404中评论按理说只要你在国内，所有的流量进出都在监控范围内，不管你怎么隐藏也没用，想搞你分..
原梓番博客在
在Laravel框架中使用模型Model分表最简单的方法中评论好久好久都没看友情链接申请了，今天刚看，已经添加。..

Top