Chapter 7 – Anonymity on the Internet
We are not anonymous on the Internet. Information is often searchable on the Internet longer than in the real world. As a result, the data needed to track a specific user will remain on the Internet for longer.
1. Introduction
Many Internet users think that when they sit at their computer in their chair in their own home, no one knows about them and that their activities on the Internet are entirely anonymous. For example, if they write a comment in a discussion forum and do not sign, no one can find out who wrote the comment. However, the opposite is true. Everything is recorded on the Internet, which is much longer durable than in the ordinary world. While in everyday life, many things are quickly forgotten and cannot be searched in memory, what gets on the Internet is indexed by search engines and can remain there for a long time. So, if there is a need to find a specific author of a text written on the Internet, it is often not such a big problem.
2. Anonymity
Anonymity is the ability to be unrecognizable, which is practically the ability to not be distinguishable from others. When moving on the Internet, we gradually leave several traces that make us less distinguishable, and therefore, the longer we move on the Internet, the less anonymous we are.
Moreover, anonymity on the Internet is more complicated than in the real world. Electronic traces persist on the Internet for a long time, are more accessible to more people, and since they are in electronic form, they can be easily searched and linked.
3. Using the Internet
Let’s first look at how the Internet works.
Connecting to the Internet
Therefore, for computers to communicate with each other, they must have unique identifiers. Similar to how we have postal addresses or social security numbers in the real world. In the world of the Internet and computer networks, they are generally called IP addresses (the abbreviation IP stands for Internet Protocol). To ensure the uniqueness of IP addresses, they are assigned hierarchically. At the top is the global organization IANA, which allocates address ranges to individual regional registrars (RIRs), then to local registrars (LIRs), and finally to particular organizations or Internet service providers (ISPs).
World division by regional registrar areas
- RIPE NCC (Réseaux IP Européens Network Coordination Centre) – Europe, Russia, Middle East and Central Asia
- APNIC (Asia Pacific Network Information Centre) – Asia, Australia and Oceania
- AFRINIC (African Network Information Centre) – Africa
- ARIN (American Registry for Internet Numbers) – Canada, USA
- LACNIC (Latin American and Caribbean Network Information Centre) – Latin America, Caribbean
Individual organizations then assign specific IP addresses to individual devices or households and keep records of who and when they assign a given IP address. This is partly to ensure the uniqueness mentioned above and to keep track of who is on their network.
In practice, this means that just by connecting to the Internet, we lose a lot of our anonymity – from the publicly available registrar database, it is possible to find out which organization the IP address was assigned to, and at the same time, the network administrators of the organization know who was assigned which IP address at what time.
What does an IP address look like?
IPv4 (Internet Protocol version 4) is a four-byte identifier written in decimal, for example, 192.168.100.10. There are about 4 billion IPv4 addresses in total.
IPv6 (Internet Protocol version 6) was created to replace IPv4 because, strangely enough, 4 billion addresses are not sufficient. Therefore, it is a sixteen-byte identifier and is written in hexadecimal to shorten the length of the writing, for example, 2001:db8:bee:b00:10ab. There are 3.4×1038 IPv6 addresses, which could be enough for the next few centuries.
Communication on the Internet
Computer network administrators also need to know what is happening on their network to be able to detect and eliminate technical problems on time. If necessary, they should also investigate cyber attacks and report security incidents. That is why data about network traffic is collected and stored – specifically, which IP address communicated with which IP address, how much data was transferred and when, etc. However, the actual transferred data is not stored because there is an enormous amount of it. It is similar to someone collecting information about who you spoke to on the phone but not storing the call’s content. Every network administrator stores this traffic data – the public connection provider (ISP) must do so by law, and others do it too because it is in their interest.
We must not forget that the Internet comprises many interconnected networks, and our communication – for example, when visiting a website on a server overseas – goes through several different networks of different organizations. Each of them collects and stores data about the traffic on its network. So, every communication leaves traces, thanks to which it is theoretically possible to assign the IP address of its originator to each one. As shown above, using an IP address is not entirely anonymous, so communication cannot be completely anonymous in principle.
Anonymization networks
Experienced users can use anonymization networks, such as TOR or I2P, to anonymize their activities. Communication here is conducted through many intermediaries, so the information about who is communicating with whom is blurred.
As the cases of tracking down several experienced cybercriminals show, one cannot rely on 100% anonymity here either. Most users of anonymization networks rely on traffic information being stored in individual networks for only a limited time, so they leave traces behind. Still, before investigators can get to them, they are overwritten by newer records.
Using Internet Services
People use the Internet for its many services, such as websites. Each such service has its operator, which stores operational information for the same reasons as the network administrator. Typically, this includes the time of visit/login and the device’s IP address from which the user accesses. It also records the time of various user actions, such as registration, posting, and editing of the profile.
User profiling is also usually performed, i.e. when they typically log in, how long they are present, what links they click on, and what interests them. This information, which the service operator can collect, is valuable for advertising vendors and other companies so that it can be the subject of trade. To know how the operator can handle this data, you need to read the service’s Terms of Use/License Terms.
What does your browser reveal about you?
When you visit a website, your web browser sends a lot of information about itself so that it can be formatted correctly (screen resolution, language settings), but also many others that can be used to reduce the anonymity of the visitor (for example, which page we come from, geolocation data, etc.).
Do you want to know what your browser reveals about itself? Take a look at https://amiunique.org
Using Internet services leaves digital traces that reduce the anonymity of the user. The more services a user uses, or the longer they use, the less anonymous they are. In conjunction with information from the network operator and the essentially non-anonymous allocation of IP addresses, anonymity on the Internet cannot be relied upon.
Pseudonymity
The term pseudonymity is also often used in connection with service providers. The information available to the operator is usually insufficient to identify a physical user fully. Still, it can be determined from it that it is the same user (or browser) as the last time. For example, online shops can offer the same type of goods the user viewed during the previous visit.
4. Disclosure of information
However, the digital traces that remain after us due to using the Internet and that we cannot influence in any way are not the only clues to our identification. Some users behave like reckless exhibitionists, and social networks provide them with space and opportunity. They often like to show off and voluntarily publish information, thereby gradually revealing their privacy. We need to be careful about what we publish about ourselves and who can see it, and realize that individual innocent fragments can be put together so that they can provide utterly unexpected information. For example, an experienced psychologist can describe your personality well based on the information you publish about yourself. HR professionals can also use this information when screening job applicants.
5. Who can identify me?
Fortunately, not everyone has access to information that fully identifies us. However, it is essential to remember that identification is possible if identification is needed, for example, during an investigation by law enforcement agencies. Therefore, it is necessary to adapt our behaviour on the Internet to this fact.
6. Summary
Anonymity is the ability not to be distinguished from others. Using the Internet, the anonymity of individual users is gradually reduced because they leave digital traces behind.
The principle of the Internet cannot prevent this because each computer must use hierarchically assigned IP addresses for communication, and it is possible to find out who they were assigned to at a particular moment. Further information about Internet usage is obtained from the service providers, such as websites. In addition to these automatically and technically created digital traces, users voluntarily publish other information about themselves on social networks, in comments, discussions, or on their websites. Alternatively, employers may also publish some information about their employees. It is, therefore a good habit not to rely on anonymity on the Internet and to behave as if we were always fully identifiable.