Skip to content

Latest commit

 

History

History
9 lines (6 loc) · 472 Bytes

README.md

File metadata and controls

9 lines (6 loc) · 472 Bytes

sscorpus: A monolingual parallel corpus for sentence simplification

This corpus contains 492,993 aligned sentences extracted by pairing Simple English Wikipedia with English Wikipedia. These source data were downloaded in May 2016.

The form of each line in the corpus: original sentence <TAB> simple sentence <TAB> similarity score

For questions, please contact Tomoyuki Kajiwara at Tokyo Metropolitan University.