Websites contain a lot text and data. Some of them have API's for easy download of data, but many don't. In this course you will learn how to harvest various parts of webpages that don't have APIs. You will learn how to inspect a website and identify the parts of it that you want to harvest, whether it be numbers or text, and use R to transform it into a workable format.
This is called web scraping. One of the advantages is that we can harvest data and text that is spread over many pages instead of having to copy-paste each page manually. In addition, we can ensure that it is in a format that allows us to work with page's content. In this course we will harvest both numbers in tables about the demographics of University of Copenhagen Students as well as text about a political and societal issue, both past and present
We assume that you have some experience working with R, and know the tidyverse. The level required is what you would learn on one of our introductory course in "R for absolute beginners".
You are kindly asked to have R and Positron and the internet browser Firefox installed before the course. You can install R for Windows here and for Mac here. You can install Positron for Windows and Mac here.
The course will be held in English if any participant requests it. If everyone speaks Danish it can be held in Danish
This course is part of the R Summer School. It is an educational programme aimed at students at Roskilde University and the University of Copenhagen. As a participant in the summer school, you will have the opportunity to work intensively with R and thereby acquire strong data skills. It is up to you whether you enrol in one or more courses.
Please note this is an online event. The online event URL will be sent to you in an email confirming your registration.
* Required Field
Registration only possible for emails from these domains - ku.dk, @kb.dk, @itu.dk, au.dk, aau.dk, sdu.dk, @ruc.dk, e.g. xyz123@alumni.ku.dk
Registration is required. There are 29 seats available.
Cand.scient.bibl.
Dataanalytiker
Københavns Universitetsbibliotek, Forskerservice & KUB Datalab
To use this platform, the system writes one or more cookies in your browser. These cookies are not shared with any third parties. In addition, your IP address and browser information is stored in server logs and used to generate anonymized usage statistics. Your institution uses these statistics to gauge the use of library content, and the information is not shared with any third parties.