question: how could I extract a specific number of keywords instead of sentence? #180

Archkik · 2022-07-21T10:12:46Z

how could I extract a specific number of keywords instead of sentence with python API?

miso-belica · 2022-07-21T12:00:38Z

You can pick from the summary anything you want by providing custom function. The function gets collection if SentenceInfo objects.

# -*- coding: utf-8 -*-

from sumy.parsers.html import HtmlParser
from sumy.nlp.tokenizers import Tokenizer
from sumy.summarizers.lsa import LsaSummarizer as Summarizer
from sumy.nlp.stemmers import Stemmer
from sumy.utils import get_stop_words


LANGUAGE = "english"


def pick_sentences(infos: list[SentenceInfo]):
	# your algorithm here
	return [] # any SentenceInfo objects you want to pick


if __name__ == "__main__":
    url = "https://en.wikipedia.org/wiki/Automatic_summarization"
    parser = HtmlParser.from_url(url, Tokenizer(LANGUAGE))
    stemmer = Stemmer(LANGUAGE)

    summarizer = Summarizer(stemmer)
    summarizer.stop_words = get_stop_words(LANGUAGE)

    for sentence in summarizer(parser.document, pick_sentences):
        print(sentence)

miso-belica · 2022-10-23T16:53:53Z

@Archkik does this work for your use-case? Is your issue different somehow? Can you describe what you are trying to achieve then?

miso-belica added the question label Jul 21, 2022

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

question: how could I extract a specific number of keywords instead of sentence? #180

question: how could I extract a specific number of keywords instead of sentence? #180

Archkik commented Jul 21, 2022

miso-belica commented Jul 21, 2022 •

edited

miso-belica commented Oct 23, 2022

question: how could I extract a specific number of keywords instead of sentence? #180

question: how could I extract a specific number of keywords instead of sentence? #180

Comments

Archkik commented Jul 21, 2022

miso-belica commented Jul 21, 2022 • edited

miso-belica commented Oct 23, 2022

miso-belica commented Jul 21, 2022 •

edited