Methods Introduction: Statistics with Lions – Part 2

Statistics & Methods

17. November 2017

Team statworx

After the descriptive examination of the historical data, our archaeologists take a step further and pose the following research question:

H1: The longer the lions are engaged in circus games, the higher their weight.

To answer this question, the researchers initially use a simple correlation analysis.

‍
This reveals that weight and months in the circus are positively correlated with Pearson’s Rho $rho$ = .7718 and a significance of p $le$ .001. Thus, the lions gained weight the longer they were in the circus.

pwcorr weight months, sig obs

However, the evaluating statisticians knew from the previous analysis that the weight gain did not progress continuously. Especially for lions that had been in the circus for a longer time, it could be graphically seen that their weight decreased again. Therefore, the scientists categorized the months to perform a variance analysis (ANOVA) with the new variable.

tab months_kat

The goal of ANOVA now is to see if the average weight between the newly formed groups is different, and (advantage ANOVA over correlation) in which direction the weight differs. The procedure imposes several requirements on the data, which are to be considered individually in part 2 of this series.

Prerequisites of an ANOVA

Assumption 1: The dependent variable (weight) should be metric. Weight corresponds to a ratio scale due to its zero point.

‍Assumption 2: The independent variable (months in categories) should consist of at least 3 categories. Our variable consists of a total of 6 categories and thus seems ideally suited.

‍Assumption 3: Independent observations should be present. This means that no relationship should exist between the observations in and between each group.

While the first three assumptions cannot be checked with a statistical procedure, this is the case with the now following assumptions.

‍Assumption 4: No significant outliers should exist. The verification of this assumption is done graphically with the help of boxplots. As you can see below, statistically speaking, there are none, as otherwise points would be present above or below the antennas in the representation.

graph box weight, over(months_kat, label(labsize(vsmall)))

Statistically speaking, calculating outliers is very simple: Starting from the upper and lower end of the box (75% and 25% quartile), the interquartile range (for each group) is calculated.

Formula: $IQR = x_{0.75} – x_{0.25}$

For the lions that have been in the circus for less than 6 months, this can be calculated with the Stata command:‍

centile weight if months_kat == 1, centile(25 75)

‍

So the 75% quartile in this group is 112kg, and the 25% quartile is 93kg. The IQR is thus 112kg - 93kg = 19kg.

An outlier "upwards" or "downwards" is defined such that 1.5 times the edge length must not be exceeded. This is considered the standard value in the calculation of outliers, so the end of the antennas in the boxplot represents exactly this threshold. This is calculated with:

$Outliers_{upwards} = |{frac{ Wert - Quartil_{0,75} }{ IQR } }|Outliers_{downwards} = |{frac{ Wert - Quartil_{0,25} }{ IQR } }| $

Using the example of lions in category 1 (less than 6 months in the circus), let's display the minimum and maximum weight.

su gewicht if monate_kat == 1

‍

The minimum weight of these lions is 75kg, and the maximum weight is 127kg. The proportion of the IQR is thus:

$Outliers_{upwards} = |{frac{ 127 - 112 }{ 19 } }| = 0.789 le 1.5 Outliers_{downwards} = |{frac{ 75 - 93 }{ 19 } }| = 0.947 le 1.5 $

Since the minimum and maximum in this group do not exceed the value of 1.5, no outliers are present. Unfortunately, Stata does not natively provide a calculation including the display of outliers according to this logic, but the command "extremes" can be additionally installed. The command performs the calculations just discussed immediately and would display outlier values.

ssc install extremes 
bysort months_kat: extreme weight, iqr(1.5)

‍‍

Since entering the command, grouped by the categories of months, shows no values, the graphical consideration of the boxplots is naturally confirmed.

‍Assumption 5: The dependent variable (weight) should be normally distributed within each category (months). Normal distribution is often considered a casus belli for the further course of the analysis. Almost stoically, the rather not recommended Kolmogorov-Smirnov test (Field 184 ff.) is often used exclusively for this purpose, without also looking at the graphical distribution itself. The histograms below consider the distribution of weight across all 6 categories of months and also give the result of the recommended Shapiro-Wilk test for normal distribution. If the significance p $le$.05 in this test (similar to Kolmogorov-Smirnov), the variable can be considered normally distributed.

For three of the six categories, the test with p $le$.05 indicates that there is no normal distribution. Visually, however, no gross violation can be detected in these 3 categories either, so there is no objection to the evaluation with ANOVA here. Furthermore, reference is made to the numerous literature on the robustness of ANOVA in violation of this assumption.

‍Assumption 6: The last assumption of ANOVA is that the variances between groups concerning the dependent variable are equal. This is also known as homogeneity of variances. With a sole focus on variance, this can hardly be assessed (see lower table), so a test for variance homogeneity must be performed.

tabstat weight, statistics(mean var) by(months_kat)

This test is the Levene test, which proves variance homogeneity with p $ge$ .05 and is requested in Stata via‍

robvar weight, by(months_kat)

The output of the Levene test is marked in Stata with W0 and shows p $le$.001 in this case, so homogeneity of variances cannot be spoken of. In this case, a correction per Welch or Brown-Forsythe should be made for the calculation of the F-values of ANOVA, and posthoc tests for the violation of this assumption also exist. Unlike SPSS, Stata does not have an option within the command to apply Welch or Brown-Forsythe, but reference is also made to the robustness of ANOVA here. Additionally, Stata, with the help of the regression command, which can be used complementary to ANOVA, has far greater possibilities of calculation when this problem arises. For our analysts, the violation of variance homogeneity is currently not an issue.

A Word

The assumptions of ANOVA are often considered sacred and immutable law. Often, when an assumption is violated, to which ANOVA reacts quite robustly, a non-parametric alternative procedure is used. Moreover, not infrequently, means over the groups are presented descriptively and then interpreted with the results (significances) of the non-parametric procedure. Overall, ANOVA is a robust procedure that can still be applied even in the face of gross violations of the assumptions.

References

Field, Andy (2013): Discovering Statistics Using SPSS, 4th Edition. SAGE: London

‍

Marcel Plaschke

Head of Strategy, Sales & Marketing

schedule a consultation

Content

Zugehörige Leistungen

More Blog Posts

Artificial Intelligence

AI Trends Report 2025: All 16 Trends at a Glance

Tarik Ashry

05. February 2025

Artificial Intelligence
Data Science
Human-centered AI

Explainable AI in practice: Finding the right method to open the Black Box

Jonas Wacker

15. November 2024

Artificial Intelligence
Data Science
GenAI

How a CustomGPT Enhances Efficiency and Creativity at hagebau

Tarik Ashry

06. November 2024

Artificial Intelligence
Data Culture
Data Science
Deep Learning
GenAI
Machine Learning

AI Trends Report 2024: statworx COO Fabian Müller Takes Stock

Tarik Ashry

05. September 2024

Artificial Intelligence
Human-centered AI
Strategy

The AI Act is here – These are the risk classes you should know

Fabian Müller

05. August 2024

Artificial Intelligence
GenAI
statworx

Back to the Future: The Story of Generative AI (Episode 4)

Tarik Ashry

31. July 2024

Artificial Intelligence
GenAI
statworx

Back to the Future: The Story of Generative AI (Episode 3)

Tarik Ashry

24. July 2024

Artificial Intelligence
GenAI
statworx

Back to the Future: The Story of Generative AI (Episode 2)

Tarik Ashry

04. July 2024

Artificial Intelligence
GenAI
statworx

Back to the Future: The Story of Generative AI (Episode 1)

Tarik Ashry

10. July 2024

Artificial Intelligence
GenAI
statworx

Generative AI as a Thinking Machine? A Media Theory Perspective

Tarik Ashry

13. June 2024

Artificial Intelligence
GenAI
statworx

Custom AI Chatbots: Combining Strong Performance and Rapid Integration

Tarik Ashry

10. April 2024

Artificial Intelligence
Data Culture
Human-centered AI

How managers can strengthen the data culture in the company

Tarik Ashry

21. February 2024

Artificial Intelligence
Data Culture
Human-centered AI

AI in the Workplace: How We Turn Skepticism into Confidence

Tarik Ashry

08. February 2024

Artificial Intelligence
Data Science
GenAI

The Future of Customer Service: Generative AI as a Success Factor

Tarik Ashry

25. October 2023

Artificial Intelligence
Data Science

How we developed a chatbot with real knowledge for Microsoft

Isabel Hermes

27. September 2023

Data Science
Data Visualization
Frontend Solution

Why Frontend Development is Useful in Data Science Applications

Jakob Gepp

30. August 2023

Artificial Intelligence
Human-centered AI
statworx

the byte - How We Built an AI-Powered Pop-Up Restaurant

Sebastian Heinz

14. June 2023

Artificial Intelligence
Recap
statworx

Big Data & AI World 2023 Recap

Team statworx

24. May 2023

Data Science
Human-centered AI
Statistics & Methods

Unlocking the Black Box – 3 Explainable AI Methods to Prepare for the AI Act

Team statworx

17. May 2023

Artificial Intelligence
Human-centered AI
Strategy

How the AI Act will change the AI industry: Everything you need to know about it now

Team statworx

11. May 2023

Artificial Intelligence
Human-centered AI
Machine Learning

Gender Representation in AI – Part 2: Automating the Generation of Gender-Neutral Versions of Face Images

Team statworx

03. May 2023

Artificial Intelligence
Data Science
Statistics & Methods

A first look into our Forecasting Recommender Tool

Team statworx

26. April 2023

Artificial Intelligence
Data Science

On Can, Do, and Want – Why Data Culture and Death Metal have a lot in common

David Schlepps

19. April 2023

Artificial Intelligence
Human-centered AI
Machine Learning

GPT-4 - A categorisation of the most important innovations

Mareike Flögel

17. March 2023

Artificial Intelligence
Data Science
Strategy

Decoding the secret of Data Culture: These factors truly influence the culture and success of businesses

Team statworx

16. March 2023

Artificial Intelligence
Deep Learning
Machine Learning

How to create AI-generated avatars using Stable Diffusion and Textual Inversion

Team statworx

08. March 2023

Artificial Intelligence
Human-centered AI
Strategy

Knowledge Management with NLP: How to easily process emails with AI

Team statworx

02. March 2023

Artificial Intelligence
Deep Learning
Machine Learning

3 specific use cases of how ChatGPT will revolutionize communication in companies

Ingo Marquart

16. February 2023

Recap
statworx

Ho ho ho – Christmas Kitchen Party

Julius Heinz

22. December 2022

Artificial Intelligence
Deep Learning
Machine Learning

Real-Time Computer Vision: Face Recognition with a Robot

Sarah Sester

30. November 2022

Data Engineering
Tutorial

Data Engineering – From Zero to Hero

Thomas Alcock

23. November 2022

Recap
statworx

statworx @ UXDX Conf 2022

Markus Berroth

18. November 2022

Artificial Intelligence
Machine Learning
Tutorial

Paradigm Shift in NLP: 5 Approaches to Write Better Prompts

Team statworx

26. October 2022

Recap
statworx

statworx @ vuejs.de Conf 2022

Jakob Gepp

14. October 2022

Data Engineering
Data Science

Application and Infrastructure Monitoring and Logging: metrics and (event) logs

Team statworx

29. September 2022

Coding
Data Science
Machine Learning

Zero-Shot Text Classification

Fabian Müller

29. September 2022

Cloud Technology
Data Engineering
Data Science

How to Get Your Data Science Project Ready for the Cloud

Alexander Broska

14. September 2022

Artificial Intelligence
Human-centered AI
Machine Learning

Gender Representation in AI – Part 1: Utilizing StyleGAN to Explore Gender Directions in Face Image Editing

Isabel Hermes

18. August 2022

Artificial Intelligence
Human-centered AI

statworx AI Principles: Why We Started Developing Our Own AI Guidelines

Team statworx

04. August 2022

Data Engineering
Data Science
Python

How to Scan Your Code and Dependencies in Python

Thomas Alcock

21. July 2022

Data Engineering
Data Science
Machine Learning

Data-Centric AI: From Model-First to Data-First AI Processes

Team statworx

13. July 2022

Artificial Intelligence
Deep Learning
Human-centered AI
Machine Learning

DALL-E 2: Why Discrimination in AI Development Cannot Be Ignored

Team statworx

28. June 2022

The helfRlein package – A collection of useful functions

Jakob Gepp

23. June 2022

Recap
statworx

Unfold 2022 in Bern – by Cleverclip

Team statworx

11. May 2022

Artificial Intelligence
Data Science
Human-centered AI
Machine Learning

Break the Bias in AI

Team statworx

08. March 2022

Artificial Intelligence
Cloud Technology
Data Science
Sustainable AI

How to Reduce the AI Carbon Footprint as a Data Scientist

Team statworx

02. February 2022

Recap
statworx

2022 and the rise of statworx next

Sebastian Heinz

06. January 2022

Recap
statworx

5 highlights from the Zurich Digital Festival 2021

Team statworx

25. November 2021

Data Science
Human-centered AI
Machine Learning
Strategy

Why Data Science and AI Initiatives Fail – A Reflection on Non-Technical Factors

Team statworx

22. September 2021

Artificial Intelligence
Data Science
Human-centered AI
Machine Learning
statworx

Column: Human and machine side by side

Sebastian Heinz

03. September 2021

Coding
Data Science
Python

How to Automatically Create Project Graphs With Call Graph

Team statworx

25. August 2021

Coding
Python
Tutorial

statworx Cheatsheets – Python Basics Cheatsheet for Data Science

Team statworx

13. August 2021

Data Science
statworx
Strategy

STATWORX meets DHBW – Data Science Real-World Use Cases

Team statworx

04. August 2021

Data Engineering
Data Science
Machine Learning

Deploy and Scale Machine Learning Models with Kubernetes

Team statworx

29. July 2021

Cloud Technology
Data Engineering
Machine Learning

3 Scenarios for Deploying Machine Learning Workflows Using MLflow

Team statworx

30. June 2021

Artificial Intelligence
Deep Learning
Machine Learning

Car Model Classification III: Explainability of Deep Learning Models With Grad-CAM

Team statworx

19. May 2021

Artificial Intelligence
Coding
Deep Learning

Car Model Classification II: Deploying TensorFlow Models in Docker Using TensorFlow Serving

12. May 2021

Coding
Deep Learning

Car Model Classification I: Transfer Learning with ResNet

Team statworx

05. May 2021

Artificial Intelligence
Deep Learning
Machine Learning

Car Model Classification IV: Integrating Deep Learning Models With Dash

Dominique Lade

05. May 2021

AI Act

Potential Not Yet Fully Tapped – A Commentary on the EU’s Proposed AI Regulation

Team statworx

28. April 2021

Artificial Intelligence
Deep Learning
statworx

Creaition – revolutionizing the design process with machine learning

Team statworx

31. March 2021

Artificial Intelligence
Data Science
Machine Learning

5 Types of Machine Learning Algorithms With Use Cases

Team statworx

24. March 2021

Recaps
statworx

2020 – A Year in Review for Me and GPT-3

Sebastian Heinz

23. Dezember 2020

Artificial Intelligence
Deep Learning
Machine Learning

5 Practical Examples of NLP Use Cases

Team statworx

12. November 2020

Data Science
Deep Learning

The 5 Most Important Use Cases for Computer Vision

Team statworx

11. November 2020

Data Science
Deep Learning

New Trends in Natural Language Processing – How NLP Becomes Suitable for the Mass-Market

Dominique Lade

29. October 2020

Data Engineering

5 Technologies That Every Data Engineer Should Know

Team statworx

22. October 2020

Artificial Intelligence
Data Science
Machine Learning

‍

Generative Adversarial Networks: How Data Can Be Generated With Neural Networks

Team statworx

10. October 2020

Coding
Data Science
Deep Learning

Fine-tuning Tesseract OCR for German Invoices

Team statworx

08. October 2020

Artificial Intelligence
Machine Learning

Whitepaper: A Maturity Model for Artificial Intelligence

Team statworx

06. October 2020

Data Engineering
Data Science
Machine Learning

How to Provide Machine Learning Models With the Help Of Docker Containers

Thomas Alcock

01. October 2020

Recap
statworx

STATWORX 2.0 – Opening of the New Headquarters in Frankfurt

Julius Heinz

24. September 2020

Machine Learning
Python
Tutorial

How to Build a Machine Learning API with Python and Flask

Team statworx

29. July 2020

Data Science
Statistics & Methods

Model Regularization – The Bayesian Way

Thomas Alcock

15. July 2020

Recap
statworx

Off To New Adventures: STATWORX Office Soft Opening

Team statworx

14. July 2020

Data Engineering
R
Tutorial

How To Dockerize ShinyApps

Team statworx

15. May 2020

Coding
Python

Making Of: A Free API For COVID-19 Data

Sebastian Heinz

01. April 2020

Frontend
Python
Tutorial

How To Build A Dashboard In Python – Plotly Dash Step-by-Step Tutorial

Alexander Blaufuss

26. March 2020

Coding
R

Why Is It Called That Way?! – Origin and Meaning of R Package Names

Team statworx

19. March 2020

Data Visualization
R

Community Detection with Louvain and Infomap

Team statworx

04. March 2020

Coding
Data Engineering
Data Science

Testing REST APIs With Newman

Team statworx

26. February 2020

Coding
Frontend
R

Dynamic UI Elements in Shiny – Part 2

Team statworx

19. Febuary 2020

Coding
Data Visualization
R

Animated Plots using ggplot and gganimate

Team statworx

14. Febuary 2020

Machine Learning

Machine Learning Goes Causal II: Meet the Random Forest’s Causal Brother

Team statworx

05. February 2020

Artificial Intelligence
Machine Learning
Statistics & Methods

Machine Learning Goes Causal I: Why Causality Matters

Team statworx

29.01.2020

Data Engineering
R
Tutorial

How To Create REST APIs With R Plumber

Stephan Emmer

23. January 2020

Recaps
statworx

statworx 2019 – A Year in Review

Sebastian Heinz

20. Dezember 2019

Artificial Intelligence
Deep Learning

Deep Learning Overview and Getting Started

Team statworx

04. December 2019

Coding
Machine Learning
R

Tuning Random Forest on Time Series Data

Team statworx

21. November 2019

Data Science
R

Combining Price Elasticities and Sales Forecastings for Sales Improvement

Team statworx

06. November 2019

Data Engineering
Python

Access your Spark Cluster from Everywhere with Apache Livy

Team statworx

30. October 2019

Recap
statworx

STATWORX on Tour: Wine, Castles & Hiking!

Team statworx

18. October 2019

Data Science
R
Statistics & Methods

Evaluating Model Performance by Building Cross-Validation from Scratch

Team statworx

02. October 2019

Data Science
Machine Learning
R

Time Series Forecasting With Random Forest

Team statworx

25. September 2019

Coding
Frontend
R

Dynamic UI Elements in Shiny – Part 1

Team statworx

11. September 2019

Machine Learning
R
Statistics & Methods

What the Mape Is FALSELY Blamed For, Its TRUE Weaknesses and BETTER Alternatives!

Team statworx

16. August 2019

Coding
Python

Web Scraping 101 in Python with Requests & BeautifulSoup

Team statworx

31. July 2019

Coding
Frontend
R

Getting Started With Flexdashboards in R

Thomas Alcock

19. July 2019

Recap
statworx

statworx summer barbecue 2019

Team statworx

21. June 2019

Data Visualization
R

Interactive Network Visualization with R

Team statworx

12. June 2019