← Back to ResearchNav home简体繁體
📖

Language & Literature

This section holds the tools you use to look at language itself rather than at a subject: corpora, and scanned classics. The CCL corpus built at Peking University splits into modern Chinese, classical Chinese, spoken and Chinese–English parallel components, so frequency and collocation questions get answered against a balanced, tagged corpus instead of whatever a general engine happens to index. Google Ngram plots how often a word or phrase appears across the digitised book corpus over roughly two centuries — useful for 'when did this term catch on' or 'which of two wordings is standard in print' — but the data is optical character recognition output, so early printing, proper nouns and spelling variants are noisy: read the trend, not the decimal. Shuge collects high-resolution scans of public-domain Chinese classics and their editions, which is as close to a first-hand source as image citation gets; most material is free, behind a registration and its own attribution terms. Before trusting any corpus here, settle three things: the time span and genres it covers, what query operators it accepts, and whether counts can be exported. Those decide how far two corpora can disagree about the same word.