{"id":175599,"date":"2026-01-28T11:24:19","date_gmt":"2026-01-28T10:24:19","guid":{"rendered":"https:\/\/liora.io\/de\/?p=175599"},"modified":"2026-08-10T13:36:47","modified_gmt":"2026-08-10T11:36:47","slug":"chi-2-mehr-ueber-diesen-unentbehrlichen-statistischen-test","status":"publish","type":"post","link":"https:\/\/liora.io\/de\/chi-2-mehr-ueber-diesen-unentbehrlichen-statistischen-test","title":{"rendered":"Chi-2 : Mehr \u00fcber diesen unentbehrlichen statistischen Test"},"content":{"rendered":"\n<p><strong>Der Chi-Quadrat-Test ist ein statistischer Test f\u00fcr Variablen, die eine endliche Anzahl von m\u00f6glichen Werten annehmen (also kategoriale Variablen). Zur Erinnerung: Ein statistischer Test ist eine Methode, um eine Hypothese, die sogenannte Nullhypothese, anzunehmen oder abzulehnen, je nachdem, wie gut sie zu den Daten passt.<\/strong><\/p>\n\n\n\n<div class=\"wp-block-buttons is-layout-flex wp-block-buttons-is-layout-flex is-content-justification-center wp-container-core-buttons-is-layout-5ee10de4\" style=\"margin-top:32px;margin-bottom:32px\"><div class=\"wp-block-button\"><a class=\"wp-block-button__link wp-element-button\" href=\"https:\/\/liora.io\/de\/weiterbildung\/data-ki\/data-scientist\">Entdecke unsere Data Scientist Weiterbildungen<\/a><\/div><\/div>\n\n\n\n<h2 class=\"wp-block-heading\" id=\"wozu-dient-der-chi-quadrat-test\">Wozu dient der Chi-Quadrat-Test?<\/h2>\n\n\n\n<p>Der Vorteil des Chi-Quadrat-Tests ist seine gro\u00dfe Bandbreite an Anwendungsm\u00f6glichkeiten:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li>Test auf \u00dcbereinstimmung mit einer a priori definierten Gesetzm\u00e4\u00dfigkeit oder einer Familie von Gesetzm\u00e4\u00dfigkeiten, z.B. Folgt die Gr\u00f6\u00dfe einer Population einer Normalverteilung? :<\/li>\n<li>Test auf Unabh\u00e4ngigkeit, Beispiel: Ist die Haarfarbe unabh\u00e4ngig vom Geschlecht?<\/li>\n<li>Test auf Homogenit\u00e4t: Sind zwei Datens\u00e4tze gleich verteilt?<\/li>\n<\/ul>\n\n\n\n<h2 class=\"wp-block-heading\" id=\"wie-funktioniert-der-test\">Wie funktioniert der Test?<\/h2>\n\n\n\n<p>Das Prinzip ist, die N\u00e4he oder Ferne zwischen der Gesetzm\u00e4\u00dfigkeit der Stichprobe und einer theoretischen Gesetzm\u00e4\u00dfigkeit mit der sogenannten Pearson-Statistik <span class=\"liora-inline-math\" role=\"math\">\u03c7<sub>Pearson<\/sub><\/span> zu vergleichen, die auf dem Chi-Quadrat-Abstand basiert.<\/p>\n\n\n\n<p>Erstes Problem: Da wir nur \u00fcber eine begrenzte Anzahl von Daten verf\u00fcgen, k\u00f6nnen wir das Gesetz der Stichprobe nicht perfekt kennen, sondern nur eine Ann\u00e4herung an dieses Gesetz, das empirische Ma\u00df.<\/p>\n\n\n\n<p>Das empirische Ma\u00df <span class=\"liora-inline-math\" role=\"math\">P\u0302<sub>n,X<\/sub><\/span> stellt die H\u00e4ufigkeit der verschiedenen beobachteten Werte dar:<\/p>\n\n\n\n<div class=\"wp-block-group has-background has-global-padding is-layout-constrained wp-container-core-group-is-layout-9089ca99 wp-block-group-is-layout-constrained\" style=\"border-radius:8px;border-left-color:#ff5c43;border-left-width:4px;background-color:#fff5f2;margin-top:24px;margin-bottom:24px;padding-top:18px;padding-right:20px;padding-bottom:18px;padding-left:20px\">\n<p class=\"has-text-align-center\"><span class=\"liora-inline-math\" role=\"math\">\u2200 x \u2208 \ud835\udd4f: P\u0302<sub>n,X<\/sub>(x) = (1 \/ n) \u00d7 \u03a3<sub>k=1<\/sub><sup>n<\/sup> 1{X<sub>k<\/sub> = x}<\/span><\/p>\n<\/div>\n\n\n\n<p><i>Formel Empirische Messung<\/i><\/p>\n\n\n\n<p>mit<\/p>\n\n\n\n<p><span class=\"liora-inline-math\" role=\"math\">X<sub>1<\/sub>, \u2026, X<sub>n<\/sub><\/span> bezeichnet die Stichprobe; <span class=\"liora-inline-math\" role=\"math\">\ud835\udd4f<\/span> ist die Menge aller m\u00f6glichen Werte.<\/p>\n\n\n\n<p>Wir definieren die Pearson-Statistik als :<\/p>\n\n\n\n<div class=\"wp-block-group has-background has-global-padding is-layout-constrained wp-container-core-group-is-layout-9089ca99 wp-block-group-is-layout-constrained\" style=\"border-radius:8px;border-left-color:#ff5c43;border-left-width:4px;background-color:#fff5f2;margin-top:24px;margin-bottom:24px;padding-top:18px;padding-right:20px;padding-bottom:18px;padding-left:20px\">\n<p class=\"has-text-align-center\"><span class=\"liora-inline-math\" role=\"math\">\u03c7<sub>Pearson<\/sub> = n \u00d7 \u03c7\u00b2(P\u0302<sub>n,X<\/sub>, P<sub>theoretisch<\/sub>) = n \u00d7 \u03a3<sub>x \u2208 \ud835\udd4f<\/sub> ((P\u0302<sub>n,X<\/sub>(x) \u2212 P<sub>theoretisch<\/sub>(x))\u00b2 \/ P<sub>theoretisch<\/sub>(x))<\/span><\/p>\n<\/div>\n\n\n\n<p><em>Statistische Formel nach Pearson<\/em><\/p>\n\n\n\n<p>Unter der Nullhypothese, d. h. dass die Stichprobenverteilung mit der theoretischen Verteilung \u00fcbereinstimmt, wird die Pearson-Statistik gegen die Chi-Quadrat-Verteilung mit d Freiheitsgraden konvergieren. Die Anzahl d der Freiheitsgrade h\u00e4ngt von der Gr\u00f6\u00dfe des Problems ab und ist im Allgemeinen die Anzahl der m\u00f6glichen Werte -1.<\/p>\n\n\n\n<div class=\"wp-block-buttons is-layout-flex wp-block-buttons-is-layout-flex is-content-justification-center wp-container-core-buttons-is-layout-5ee10de4\" style=\"margin-top:32px;margin-bottom:32px\"><div class=\"wp-block-button\"><a class=\"wp-block-button__link wp-element-button\" href=\"https:\/\/liora.io\/de\/weiterbildung\/data-ki\/data-scientist\">Data Scientist Weiterbildung entdecken<\/a><\/div><\/div>\n\n\n\n<p>Zur Erinnerung: Die Chi-Quadrat-Verteilung mit <span class=\"liora-inline-math\" role=\"math\">d<\/span> Freiheitsgraden, <span class=\"liora-inline-math\" role=\"math\">\u03c7\u00b2(d)<\/span>, ist die Verteilung der Summe der Quadrate von <span class=\"liora-inline-math\" role=\"math\">d<\/span> unabh\u00e4ngigen standardnormalverteilten Zufallsvariablen:<\/p>\n\n\n\n<div class=\"wp-block-group has-background has-global-padding is-layout-constrained wp-container-core-group-is-layout-9089ca99 wp-block-group-is-layout-constrained\" style=\"border-radius:8px;border-left-color:#ff5c43;border-left-width:4px;background-color:#fff5f2;margin-top:24px;margin-bottom:24px;padding-top:18px;padding-right:20px;padding-bottom:18px;padding-left:20px\">\n<p class=\"has-text-align-center\"><span class=\"liora-inline-math\" role=\"math\">\u03c7\u00b2(d) := \u03a3<sub>k=1<\/sub><sup>d<\/sup> X<sub>k<\/sub>\u00b2, wobei X<sub>k<\/sub> \u223c N(0,1)<\/span><\/p>\n<\/div>\n\n\n\n<p>Andernfalls wird diese Statistik ins Unendliche divergieren, was die Entfernung zwischen empirischen und theoretischen Verteilungen widerspiegelt.<\/p>\n\n\n\n<div class=\"wp-block-group has-background has-global-padding is-layout-constrained wp-container-core-group-is-layout-9089ca99 wp-block-group-is-layout-constrained\" style=\"border-radius:8px;border-left-color:#ff5c43;border-left-width:4px;background-color:#fff5f2;margin-top:24px;margin-bottom:24px;padding-top:18px;padding-right:20px;padding-bottom:18px;padding-left:20px\">\n<p class=\"has-text-align-center\"><span class=\"liora-inline-math\" role=\"math\">Unter H<sub>0<\/sub>: lim<sub>n \u2192 \u221e<\/sub> \u03c7<sub>Pearson<\/sub> = \u03c7\u00b2(d).<br>Unter H<sub>1<\/sub>: lim<sub>n \u2192 \u221e<\/sub> \u03c7<sub>Pearson<\/sub> = \u221e.<\/span><\/p>\n<\/div>\n\n\n\n<p><em>Grenzformel<\/em><\/p>\n\n\n\n<h3 class=\"wp-block-heading\" id=\"was-sind-seine-vorteile\">Was sind seine Vorteile?<\/h3>\n\n\n\n<p>Wir haben also eine einfache Entscheidungsregel: Wenn die Pearson-Statistik einen bestimmten Schwellenwert \u00fcberschreitet, lehnen wir die Ausgangshypothese (die theoretische Verteilung passt nicht zu den Daten) ab, ansonsten akzeptieren wir sie. Der Vorteil des Chi-Quadrat-Tests ist, dass dieser Schwellenwert nur von der Chi-Quadrat-Verteilung und dem Alpha-Konfidenzniveau abh\u00e4ngt, also unabh\u00e4ngig von der Verteilung der Stichprobe ist.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\" id=\"eine-anwendung-der-unabhangigkeitstest\">Eine Anwendung, der Unabh\u00e4ngigkeitstest :<\/h3>\n\n\n\n<p>Nehmen wir ein Beispiel, um diesen Test zu veranschaulichen: Wir wollen wissen, ob die Geschlechter der ersten beiden Kinder X und Y eines Paares unabh\u00e4ngig sind?<\/p>\n\n\n\n<p>Wir haben die Daten in einer Kontingenztabelle zusammengefasst:<\/p>\n\n\n\n<figure class=\"wp-block-table is-style-stripes\"><table><thead><tr><th>X \/ Y<\/th><th>Kind 2: Sohn<\/th><th>Kind 2: Tochter<\/th><th>Gesamt<\/th><\/tr><\/thead><tbody><tr><th>Kind 1: Sohn<\/th><td>857<\/td><td>801<\/td><td>1658<\/td><\/tr><tr><th>Kind 1: Tochter<\/th><td>813<\/td><td>828<\/td><td>1641<\/td><\/tr><tr><th>Gesamt<\/th><td>1670<\/td><td>1629<\/td><td>3299<\/td><\/tr><\/tbody><\/table><\/figure>\n\n\n\n<p>Die Pearson-Statistik bestimmt, ob das empirische Ma\u00df der gemeinsamen Gesetzm\u00e4\u00dfigkeit (X,Y) gleich dem Produkt der marginalen empirischen Ma\u00dfe ist, was die Unabh\u00e4ngigkeit charakterisiert:<\/p>\n\n\n\n<div class=\"wp-block-group has-background has-global-padding is-layout-constrained wp-container-core-group-is-layout-9089ca99 wp-block-group-is-layout-constrained\" style=\"border-radius:8px;border-left-color:#ff5c43;border-left-width:4px;background-color:#fff5f2;margin-top:24px;margin-bottom:24px;padding-top:18px;padding-right:20px;padding-bottom:18px;padding-left:20px\">\n<p class=\"has-text-align-center\"><span class=\"liora-inline-math\" role=\"math\">\u03c7<sub>Pearson<\/sub> = n \u00d7 \u03c7\u00b2(P\u0302<sub>X\u00d7Y<\/sub>, P\u0302<sub>X<\/sub> \u00d7 P\u0302<sub>Y<\/sub>) = \u03a3<sub>x,y \u2208 {Tochter, Sohn}<\/sub> ((Beobachtung<sub>x,y<\/sub> \u2212 Theorie<sub>x,y<\/sub>)\u00b2 \/ Theorie<sub>x,y<\/sub>)<\/span><\/p>\n<\/div>\n\n\n\n<div class=\"wp-block-buttons is-layout-flex wp-block-buttons-is-layout-flex is-content-justification-center wp-container-core-buttons-is-layout-5ee10de4\" style=\"margin-top:32px;margin-bottom:32px\"><div class=\"wp-block-button\"><a class=\"wp-block-button__link wp-element-button\" href=\"https:\/\/liora.io\/de\/weiterbildung\/data-ki\/data-scientist\">Data Scientist Weiterbildung entdecken<\/a><\/div><\/div>\n\n\n\n<p>Hier ist <span class=\"liora-inline-math\" role=\"math\">Beobachtung(x,y)<\/span> die H\u00e4ufigkeit des Wertepaares <span class=\"liora-inline-math\" role=\"math\">(x,y)<\/span>:<\/p>\n\n\n\n<div class=\"wp-block-group has-background has-global-padding is-layout-constrained wp-container-core-group-is-layout-9089ca99 wp-block-group-is-layout-constrained\" style=\"border-radius:8px;border-left-color:#ff5c43;border-left-width:4px;background-color:#fff5f2;margin-top:24px;margin-bottom:24px;padding-top:18px;padding-right:20px;padding-bottom:18px;padding-left:20px\">\n<p class=\"has-text-align-center\"><span class=\"liora-inline-math\" role=\"math\">\u2200 x,y \u2208 {Tochter, Sohn}: Beobachtung<sub>x,y<\/sub> = (1 \/ n) \u00d7 \u03a3<sub>k=1<\/sub><sup>n<\/sup> 1{(X<sub>k<\/sub>,Y<sub>k<\/sub>) = (x,y)}<\/span><\/p>\n<\/div>\n\n\n\n<p>Zum Beispiel:<\/p>\n\n\n\n<div class=\"wp-block-group has-background has-global-padding is-layout-constrained wp-container-core-group-is-layout-9089ca99 wp-block-group-is-layout-constrained\" style=\"border-radius:8px;border-left-color:#ff5c43;border-left-width:4px;background-color:#fff5f2;margin-top:24px;margin-bottom:24px;padding-top:18px;padding-right:20px;padding-bottom:18px;padding-left:20px\">\n<p class=\"has-text-align-center\"><span class=\"liora-inline-math\" role=\"math\">Beobachtung(Tochter, Tochter) = 828 \/ 3299 \u2248 0,251<\/span><\/p>\n<\/div>\n\n\n\n<p>F\u00fcr <span class=\"liora-inline-math\" role=\"math\">Theorie(x,y)<\/span> wird angenommen, dass X und Y unabh\u00e4ngig sind. Die theoretische Verteilung entspricht daher dem Produkt der Randverteilungen:<\/p>\n\n\n\n<div class=\"wp-block-group has-background has-global-padding is-layout-constrained wp-container-core-group-is-layout-9089ca99 wp-block-group-is-layout-constrained\" style=\"border-radius:8px;border-left-color:#ff5c43;border-left-width:4px;background-color:#fff5f2;margin-top:24px;margin-bottom:24px;padding-top:18px;padding-right:20px;padding-bottom:18px;padding-left:20px\">\n<p class=\"has-text-align-center\"><span class=\"liora-inline-math\" role=\"math\">\u2200 x,y \u2208 {Tochter, Sohn}: Theorie<sub>x,y<\/sub> = Beobachtung<sup>X<\/sup><sub>x<\/sub> \u00d7 Beobachtung<sup>Y<\/sup><sub>y<\/sub><\/span><\/p>\n<\/div>\n\n\n\n<p>Die theoretische Wahrscheinlichkeit f\u00fcr (Sohn,Sohn) ist also:<\/p>\n\n\n\n<div class=\"wp-block-group has-background has-global-padding is-layout-constrained wp-container-core-group-is-layout-9089ca99 wp-block-group-is-layout-constrained\" style=\"border-radius:8px;border-left-color:#ff5c43;border-left-width:4px;background-color:#fff5f2;margin-top:24px;margin-bottom:24px;padding-top:18px;padding-right:20px;padding-bottom:18px;padding-left:20px\">\n<p class=\"has-text-align-center\"><span class=\"liora-inline-math\" role=\"math\">Theorie(Sohn, Sohn) = ((857 + 801) \/ 3299) \u00d7 ((857 + 813) \/ 3299) = (1658 \u00d7 1670) \/ 3299\u00b2 \u2248 0,254<\/span><\/p>\n<\/div>\n\n\n\n<p>Berechnen wir die Teststatistik mithilfe des folgenden Python-Codes:<\/p>\n\n\n\n<p>In unserem Fall haben die Variablen X und Y nur zwei m\u00f6gliche Werte: M\u00e4dchen oder Jungen. Die Dimension des Problems ist also (2-1)(2-1) oder 1.<\/p>\n\n\n\n<p>Wir vergleichen daher die Teststatistik mit dem Chi-Quantil bei 1 Freiheitsgrad \u00fcber die Funktion chi2.ppf in scipy.stats. Sie ist kleiner als das Quantil und der p-Wert ist gr\u00f6\u00dfer als das Konfidenzniveau = 0,05. Wir k\u00f6nnen die Nullhypothese mit 95%igem Vertrauen nicht ablehnen und schlie\u00dfen daher auf die Unabh\u00e4ngigkeit des Geschlechts der ersten beiden Kinder.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\" id=\"wo-liegen-seine-grenzen\">Wo liegen seine Grenzen?<\/h3>\n\n\n\n<p>Der <strong>Chi-Quadrat-Test<\/strong> scheint sehr praktisch zu sein, hat aber auch seine Grenzen: Er stellt nur fest, dass es Korrelationen gibt, aber er erkennt weder die St\u00e4rke dieser Korrelationen noch Kausalit\u00e4ten.<\/p>\n\n\n\n<div class=\"wp-block-buttons is-layout-flex wp-block-buttons-is-layout-flex is-content-justification-center wp-container-core-buttons-is-layout-5ee10de4\" style=\"margin-top:32px;margin-bottom:32px\"><div class=\"wp-block-button\"><a class=\"wp-block-button__link wp-element-button\" href=\"https:\/\/liora.io\/de\/weiterbildung\/data-ki\/data-scientist\">Data Scientist Weiterbildung entdecken<\/a><\/div><\/div>\n\n\n\n<p>Er beruht auf der Ann\u00e4herung des Chi-Quadrat-Gesetzes durch die Pearson-Statistik, die nur dann \u00fcberpr\u00fcft werden kann, wenn eine ausreichende Anzahl von Daten vorliegt. In der Praxis sieht diese G\u00fcltigkeitsbedingung wie folgt aus:<\/p>\n\n\n\n<div class=\"wp-block-group has-background has-global-padding is-layout-constrained wp-container-core-group-is-layout-9089ca99 wp-block-group-is-layout-constrained\" style=\"border-radius:8px;border-left-color:#ff5c43;border-left-width:4px;background-color:#fff5f2;margin-top:24px;margin-bottom:24px;padding-top:18px;padding-right:20px;padding-bottom:18px;padding-left:20px\">\n<p class=\"has-text-align-center\"><span class=\"liora-inline-math\" role=\"math\">\u2200 x \u2208 \ud835\udd4f: n \u00d7 P<sub>theoretisch<\/sub>(x) \u00d7 (1 \u2212 P<sub>theoretisch<\/sub>(x)) \u2265 5<\/span><\/p>\n<\/div>\n\n\n\n<p>Der exakte Test nach Fisher kann diesen Mangel beheben, erfordert aber eine hohe Rechenleistung (in der Praxis wird er auf 2*2-Kontingenztabellen beschr\u00e4nkt).<\/p>\n\n\n\n<p>Statistische Tests sind in der Data Science unerl\u00e4sslich, um die Relevanz der erkl\u00e4renden Variablen zu \u00fcberpr\u00fcfen und die Hypothesen der Modellierung zu validieren. Weitere Informationen \u00fcber Chi-2 und andere statistische Tests findest du in unserem Modul 104 ,  Explorative Statistik.<\/p>\n\n\n\n\n\n<h2 class=\"wp-block-heading\" id=\"referenzen\">Referenzen:<\/h2>\n\n\n\n<ul class=\"wp-block-list\">\n<li><a href=\"https:\/\/docs.scipy.org\/doc\/scipy\/reference\/generated\/scipy.stats.chi2.html\">SciPy-Dokumentation: Chi-Quadrat-Verteilung<\/a><\/li>\n\n\n<li><a href=\"https:\/\/docs.scipy.org\/doc\/scipy\/reference\/generated\/scipy.stats.chi2_contingency.html\">SciPy-Dokumentation: Chi-Quadrat-Kontingenztest<\/a><\/li>\n<\/ul>\n\n","protected":false},"excerpt":{"rendered":"<p>Der Chi-Quadrat-Test ist ein statistischer Test f\u00fcr Variablen, die eine endliche Anzahl von m\u00f6glichen Werten annehmen (also kategoriale Variablen). Zur Erinnerung: Ein statistischer Test ist eine Methode, um eine Hypothese, die sogenannte Nullhypothese, anzunehmen oder abzulehnen, je nachdem, wie gut sie zu den Daten passt.<\/p>\n","protected":false},"author":78,"featured_media":175600,"comment_status":"open","ping_status":"open","sticky":false,"template":"elementor_theme","format":"standard","meta":{"_acf_changed":false,"editor_notices":[],"footnotes":""},"categories":[2472],"class_list":["post-175599","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-data-ki"],"acf":[],"_links":{"self":[{"href":"https:\/\/liora.io\/de\/wp-json\/wp\/v2\/posts\/175599","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/liora.io\/de\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/liora.io\/de\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/liora.io\/de\/wp-json\/wp\/v2\/users\/78"}],"replies":[{"embeddable":true,"href":"https:\/\/liora.io\/de\/wp-json\/wp\/v2\/comments?post=175599"}],"version-history":[{"count":5,"href":"https:\/\/liora.io\/de\/wp-json\/wp\/v2\/posts\/175599\/revisions"}],"predecessor-version":[{"id":225470,"href":"https:\/\/liora.io\/de\/wp-json\/wp\/v2\/posts\/175599\/revisions\/225470"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/liora.io\/de\/wp-json\/wp\/v2\/media\/175600"}],"wp:attachment":[{"href":"https:\/\/liora.io\/de\/wp-json\/wp\/v2\/media?parent=175599"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/liora.io\/de\/wp-json\/wp\/v2\/categories?post=175599"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}