“Theorizing from Data: Avoiding the Capital Mistake
Peter Norvig
“”It is a capital mistake to theorize before one has data.”” Sir Arthur Conan Doyle’s words from 1891 remain true today. Researchers in computational linguistics and information retrieval now have a million times more data than was available 30 years ago. This talk explores what this data can do for problems in language understanding, translation, information extraction, and inference, and extrapolates to what more data may bring in the future. ”



