> For the complete documentation index, see [llms.txt](https://concordia-infant-research-lab.gitbook.io/lab-wiki/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://concordia-infant-research-lab.gitbook.io/lab-wiki/project-workflow-data-analysis/hypothesis-testing-for-categorical-variables.md).

# Hypothesis testing for categorical variables

There are several methods to test hypotheses using categorical data. The most useful methods are to calculate **confidence intervals** for proportions or for difference in proportions and to use **hypothesis tests** for proportions, difference in proportions, independence between variables and goodness of fit. Below we will briefly explore these methods.

### Confidence intervals&#x20;

While there are many summary and descriptive statistics that can be calculated for categorical data, there is no straightforward way to calculate confidence intervals. Below we explore several ways to approximate confidence intervals.

**Bootstrapping** is a technique that samples with replacement from the data and performs the same calculation through several iterations. For example, if a proportion is calculated 100 times from different samplings of the data provided, we can approximate that the true proportion of the category in the data is contained within the lower and upper limit of the values calculated in those 100 iterations.&#x20;

**Helper functions** from the **infer library** that allow you to calculate confidence intervals through bootstrapping:

`install.packages ="infer"`

<mark style="color:blue;">**Specify**</mark> allows you to specify the variable of interest in the data and the value for which to calculate the summary statistic. For example the value "female" from the variable "gender".

<mark style="color:blue;">**Generate**</mark> allows you to perform the bootstrapping operation, and specify the number of iterations for the operation.

<mark style="color:blue;">**Calculate**</mark> allows you to specify the summary statistic to be calculated for the variable of interest throughout the bootstrapping operation.

Example:

`CI <- df%>% specify(response = gender, success="female")%>%`

&#x20;           `generate(reps=100, type="bootstrap)%>%`

&#x20;           `calculate(stat="prop")`

**Phat standard distribution** it is possible to calculate the standard error (SE) from the phat standard distribution. Once the standard error is known, a confidence interval can be calculated. The SE of phat is calculated as `sqrt(phat*(1-phat)/n).` Note that for this method to be used it is necessary that the observations are independent from each other and that the n is large enough (at least n>10).

Here is an example taken from [this page.](https://rstudio-pubs-static.s3.amazonaws.com/154137_6fe804f283e44a598e90a9f6a01ea7a8.html)

<figure><img src="https://3390241049-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FG4X3PM9X7Cp5gi9bcjVG%2Fuploads%2F5cliJcN0O49h3z0WHUZE%2Fimage.png?alt=media&amp;token=a45f9de7-c7fd-481a-bcd9-e3a2b9df29d4" alt=""><figcaption></figcaption></figure>

### Hypothesis testing for a proportion

As reviewed in the section above, the r package **infer** allows calculating confidence intervals through bootstrapping operations using the helper functions **specify**, **generate** and **calculate**. To test a hypothesis we add the helper function:

<mark style="color:blue;">**Hypothesize**</mark> allows you to declare a null hypothesis. For more documentation go [here](https://infer.tidymodels.org/reference/hypothesize.html).&#x20;

Example:

`Object_name <- df%>% specify(response = gender, success="female")%>%`

&#x20;                    `hypothesize(null="point", p=0.5),`           &#x20;

&#x20;                    `generate(reps=100, type="simulate")%>%`

&#x20;                    `calculate(stat="prop")`

This method allows you to calculate the proportion value you would get under the null hypothesis. To calculate a **p-value** from there you would need to calculate the amount of phats that were more extreme than the observed phat under the null hypothesis. For example:

`null %>% summarize(mean(stat>phat)%>%`

&#x20;        `pull()*2`

Note that we multiply the pull value times 2 to account for the left tail.

### Creating contingency tables

You might want to calculate the association between categorical variables under the null hypothesis that the variables of interest are independent from each other. One way to test this hypothesis is to create **contingency tables** and to use the **chisquare statistic** to reject the null.&#x20;

To create tables you will need to use the "broom" r package: `install.packages("broom").`

Some **helper functions** when using the broom package to do contingency tables are:

<mark style="color:blue;">**Table**</mark> allows you to turn a selection of columns into a table object.

<mark style="color:blue;">**Tidy**</mark> allows you to automatically organize a table object in a more efficient way.

<mark style="color:blue;">**Uncount(n)**</mark> allows you to get a row per observation instead of a summary of the counts.

Once the variables of interest are selected and saved into a tidy table object, we can test the **independence of the variables**. To test for indenpendence we go back to our helper functions **specify, hypothesize, generate** and **calculate**. This time we generate repetitions by permutation.

Example:

`null<- data %>% specify (var1~var2) %>%`

&#x20;               `hypothesize (null="independence")%>%`

&#x20;               `generate(reps=100, type="permute")%>%`

&#x20;               `calculate(stat="Chisq")`

This allows us to generate expected counts under the null hypothesis of independence, and thus to compare them to our observed counts.

### Chisquare as a goodness of fit measure

A goodness of fit test tries to establish weather the observed values match expected values under a specific model. For example, say we expect the distribution of family language strategy use to be equal in a group of randomly sampled parents. If we have three possible family language strategies we expect this model to be true: `model<- c(opol=1/3, 2_bilingual=1/3, 1_bilingual=1/3)`.&#x20;

Once we have a model of how we expect the data to behave, we can use the **chi-squared tes**t as a **goodness-of-fit measure,** by comparing our model to our observed data a such:

`chisq.test(observed_table, p=model)$stat`
