Latent Dirichlet Allocation (LDA) based topic modeling of news corpus in Apache Spark Streaming
- Java 1.8.x
- Scala 2.10.x
- Spark 1.6 +
- Any OS
- 6GB+ RAM
AG is a collection of more than 1 million news articles. News articles have been gathered from more than 2000 news sources by ComeToMyHead in more than 1 year of activity. ComeToMyHead is an academic news search engine which has been running since July, 2004.
The dataset is provided by the academic comunity for research purposes in data mining (clustering, classification, etc), information retrieval (ranking, search, etc), xml, data compression, data streaming, and any other non - commercial activity.
The code will automatically download the yahoo news corpus Ref . In case of some issue, you can directly download the news corpus file (118 MB) from here