الانتقال إلى المحتوى الرئيسي
Learning LoftInstitute

الرياضيات

Outliers

A value far from the rest. The 1.5 × IQR rule applied to real data, what one outlier does to the mean, and why you investigate it instead of deleting it.

Outlier

An outlier is a data value that lies far enough away from the rest of the data set to stand apart from it.

يُسمّى أيضاً
Anomaly, Extreme value
أين يقابله الطلاب
Grade 8 informally, when a scatter graph has one stray point, and formally with the 1.5 × IQR rule in GCSE Statistics and at A level.

الإجابة باختصار

An outlier is a value far away from the rest of a data set. The usual test is the 1.5 × IQR rule: anything below Q1 − 1.5 × IQR, or above Q3 + 1.5 × IQR, counts as an outlier. Finding one is a reason to investigate the value, not a licence to delete it.

مثال

Q1 = 25, Q3 = 45, IQR = 20 → boundaries at 25 − 30 = −5 and 45 + 30 = 75; 130 is an outlier

Eleven students report their daily study minutes: 20, 25, 25, 30, 30, 35, 35, 40, 45, 50, 130. The quartiles sit at the 3rd and 9th values, so Q1 = 25 and Q3 = 45, giving an IQR of 20. One and a half times that is 30, so anything below −5 or above 75 is flagged. Only 130 qualifies, and it is worth asking whether that student really studies over two hours a day or whether 13.0 was typed as 130.

A rule instead of an opinion

"Far from the rest" is a matter of taste until it is given a boundary, and different people looking at the same graph will point at different dots. The 1.5 × IQR rule fixes that by measuring distance in units of the data's own spread: a value is an outlier if it sits more than one and a half interquartile ranges outside the box.

That relative measure is the clever part. In a tightly grouped data set the boundaries close in and a modest deviation is flagged; in a widely spread one the boundaries open out and the same absolute gap is not. The rule adapts to the data rather than imposing a fixed distance, which is why it survives being applied to test marks, rainfall and reaction times alike.

What one outlier does to the summary

The eleven study times total 465, so the mean is 42.3 minutes. Remove the 130 and the remaining ten total 335, giving a mean of 33.5. One value out of eleven moved the mean by nearly nine minutes, and the mean it produced was higher than nine of the eleven actual values.

The median barely notices. It goes from 35 to 32.5 across the same change, because it depends on position rather than size — the outlier is still just "the top one" whether it reads 130 or 1,300. The same asymmetry runs through the measures of spread: the range collapses from 110 to 30 when the outlier goes, while the interquartile range moves hardly at all.

What to actually do with one

Investigate before you decide. Outliers come from three quite different places, and they deserve different treatment: a recording error, a genuine but unusual case, or a value that belongs to a different population altogether. A 130 that turns out to be a mis-typed 13.0 should be corrected. A student who genuinely studies 130 minutes a day should stay in.

Deleting a value because it spoils a pattern is how you get a wrong answer that looks tidy, and it is the habit that turns a statistics exercise into a bad one. If you do exclude a point, say so and say why, and quote the summary both with and without it. On an exam paper the mark is usually for identifying the outlier and commenting on its effect, not for removing it.

أسئلة شائعة

Does an outlier always have to be removed?

No, and removing one by default is bad practice. A genuine extreme value is data, and it may be the most interesting thing in the set — the one faulty component, the one exceptional result. Remove a value only when you can identify it as an error, and state that you have done so alongside the reason.

Where does 1.5 come from in the 1.5 × IQR rule?

It is a convention rather than a derived constant, chosen because it flags genuinely unusual values without flagging too many ordinary ones. Some analysts use 3 × IQR for a stricter "extreme outlier" boundary. Because it is a convention, exam questions state the rule they want you to use rather than assuming.

Can there be more than one outlier?

Certainly, at either end or both. Apply the rule and every value outside the boundaries is flagged. If a large fraction of the data is being flagged, that is usually a sign the data is skewed rather than that it is full of errors, and the rule is telling you about the shape of the distribution.

How do I spot an outlier on a scatter graph?

Look for the point that sits away from the pattern the others form, not simply the point furthest from the origin. A point can have an ordinary x-value and an ordinary y-value yet still be an outlier because that combination breaks the trend. On correlation questions, examiners usually expect it to be ignored when drawing the line of best fit and mentioned in the comment.

آخر تحديث

معرفة الكلمة ليست كاستخدامها

يستطيع المعلّم أن يرى الطالب وهو يستعمله في سؤال، فيتبيّن أين يتوقف الفهم بالضبط. الحصة الأولى مجانية.

اختياري
اختياري
المواد

اختر كل ما تريد تغطيته

نوع الحصة
اختياري
اختياري

كلما كنت أكثر تحديدًا، كان اختيارنا للمعلم أدق.

راسلنا على WhatsApp