Six Types of AI Bias Everybody Should Know

Author: Ed Shee Fairness. In my last blog we discussed the differences between Bias and Fairness in AI. Although I gave a basic overview of Bias, we will now go in greater detail. Bias is a component of machine learning, in many different ways. It is important to remember that raising a machine-learning model is very similar to parenting a child. MI PHAM, Unsplash Photo: A child learns from their environment using senses such as hearing, sight, and touch. Your upbringing will influence your understanding of the world and how you make decisions. A child who grows up in a heterosexual community might not realize that they are biased towards different genders. The machine learning models work exactly the same. They use data instead of sensing inputs. Data that they *weprovide! It is important that biases in data for machine learning models are avoided. We’ll take a look at the most prevalent forms of machine learning bias: Historical bias. When gathering data to train a machine-learning algorithm, it is often the easiest way to begin. It’s easy to add bias in historical data if we don’t pay attention. Amazon is an example. In 2014, they created a system to automatically screen job candidates. It was simple to feed hundreds of CVs into the system and then have top applicants selected automatically. It was trained using 10 years of job applications, and the results. Problem? The problem? Because there are more women than men at Amazon, the algorithm discovered that men were better candidates. It actively discriminated against male applicants. 2015 The entire project was scrapped. Sample Bias: When your data is not representative of real-world usage, it’s called sample bias. One population will be either overrepresented or underrepresented. David Keene gave an excellent example of sample bias in a recent talk. You need lots of audio clips and their transcriptions to train a speech-to text system. Audiobooks are a great way to gather this information. This approach is so simple. It turns out, the majority of audiobooks are read by middle-aged, educated white men. This approach results in speech recognition software that is less effective when the user comes from different socio-economic backgrounds or ethnicities. Source: https://www.pnas.org/content/117/14/7684 The chart above shows the word error rate [WER] for speech recognition systems from big tech companies. It is clear that algorithms perform poorly for voices of color compared to those with white voices. Before ML algorithms can be used, a lot of data must first be labeled. This is something you can do yourself when you log into websites. Have you ever been asked to find the traffic light squares? In reality, you’re confirming the labeling of an image in order to train visual recognition algorithms. However, the way we label data varies greatly and biases in labelling could cause problems. Image by Author You can train a system to label lions with the boxes in the above images. Then, you show the system the following image. Image by Author It is not able to recognize the obvious lion in this picture. You have inadvertently biased the system toward front-facing pictures of lions by labeling only faces. Aggregation Bias We sometimes combine data in order to make it easier or show it in a certain way. This could lead to bias, regardless of when it happened before or after we created our model. This chart shows salary growth based upon years of experience in the job. This shows that the greater your work experience, the higher the salary you will receive. Now let’s look at this data: For athletes, the exact opposite is true. While they can earn high salaries while still being at peak physical performance, it drops as soon as they cease competing. We are making the algorithm bias against these people by combining them with others in other fields. Confirmation Bias Simply stated, confirmation bias refers to our tendency not to believe information that supports our beliefs and discard those that don’t. Although I can theoretically build the best ML system, with no bias in the data nor the modelling, if you are going to alter the result based upon your “gut feeling”, it won’t really matter. Machine learning applications are particularly susceptible to confirmation bias. Human review of any decision is necessary before it is made. Doctors have been dismissive about algorithmic diagnoses in healthcare because they don’t agree with their experience. When doctors are questioned, they often discover that they haven’t reviewed the latest research, which can lead to different diagnosis outcomes, symptoms and techniques. There are only so many journals one doctor can study (especially if they’re saving people’s lives), but an ML system is able to read all of them. Photo of Evaluation Bias by Arnaud Jacobs on Unsplash. Let’s say you are building a machine-learning model that can predict the turnout in a country’s general election. By taking into account a number of factors such as age, occupation, income, and political alignment you hope to accurately predict who will vote. Your model is built, you test it in a local election, and your results are amazing. You can predict whether someone will vote most of the times. You are shocked when the general election arrives. Your model, which you had spent hours designing and testing, was not correct 55% all the time. It performed marginally worse than random guesses. Evaluation bias is evident in the poor results. You have accidentally created a system that works only for people who live in the same area as your model. Even though they are included in the initial training data, other areas with completely different voting patterns have not been adequately accounted for. We’ve seen six ways bias could impact machine learning. Although it is not an exhaustive list it will give you a solid understanding of how ML systems can become biased. Mehrabi and colleagues have a paper that I recommend if you are interested. Six Types of AI Bias Everybody Should Know originally appeared in on Medium. People are responding and highlighting this story. Published by Instant vortex Plus 10 Quart Air Fryer and Rotisserie, Convection Ovens, Roast, Bake and Dehydrate, 1500W. Also available for Amazon Prime $139. 99 $89. 95 As of December 4, 2021

THE FOREFRONT OF TECHNOLOGY

We monitors and writes about new technologies in areas such as technology, innovation, digitization, space, Earth, IT and AI.

Related Posts

Leave a Reply