ANALYSIS

Crosstabs

If you click on the 'ANALYSE' - 'Statistics' tab and select 'Crosstabs' from the drop-down menu ('Select:'), you have the option of setting up a two-dimensional table of two variables where you can look at how they are correlated.

In the first drop-down menu you select the independent variable and in the second drop-down menu you select the dependent variable. The system generates a correlation table with frequencies and colored boxes according to the correlation. To the right of the selected variables you can also select 'residuals'. The legend of the correlation according to the residuals is shown on the left below the table. If you want the results in percentages, then select the 'Percentage' option above. You can also choose to display graphs for the table by turning on the 'Show graphs' option. The left-hand side above the table displays the Chi-square coefficient.

2KA and 3KA package users as well as users of the business packages for groups can export the Crosstabs table to PDF, MS Word or MS Excel formats. These users can also analyse statistics with advanced analysis options and filter them according to segments, zoom, groups, variables, conditions, missing values, time period and unit status.

The analysis settings are located in the upper right corner in the form of a 'wheel' icon. Here you can set:

  • Basic analysis view:
    • Among other things, you can specify which types of questions and categories are displayed (categories, category other, numbers, text questions) - all are displayed by default.
  • Basic analysis settings:
    • Decimal settings.
  • Chart settings:
    • Alignment,
    • Font size,
    • Position of numerus,
    • Display average on the chart.

Some of the functionalities described are only available for users of the 2KA and 3KA packages for individuals, and for users of the business packages for groups.


Related content

Limits of Displaying Answers in Analyses

In case of a large number of answers, only 30 answers are displayed by default in the ANALYSE' - 'Statistics' tab, which is noticeable for variables with open (text) responses.

The default setting of 30 answers can be changed in the 'Number of open answers' drop-down menu in the top right.

The setting can also be changed by clicking on the 'Settings' icon, then selecting the desired value in the  'Default number of answers' drop-down menu in the 'Basic analysis settings' section and clicking on the 'Run' button.

 

Related content

Residuals in Crosstabs

When a crosstab analysis is created in the 'ANALYSE' – 'Statistics' – 'Crosstabs' tab, the value of the Chi-square is displayed and the cells within the table are coloured, based on residuals.

Residuals make it extremely easy and efficient to analyse what is happening in the table. Unlike the Chi-square, which gives only a general diagnosis of the relationship in the table, residuals show exactly where the correlation is happening. In fact, a Chi-square may be statistically significant only because of a correlation in a single cell, but it does not tell us where that is.

Residual is a term from the analysis of nominal variables. A residual is simply the difference between the actual frequency in a given cell and the theoretical frequency that would exist if the variables of a two-dimensional table in that cell were uncorrelated (the null hypothesis). The theoretical frequency is calculated very simply as the product of the two margins divided by the total size of the table.

If the underlying residuals - which follow a Poisson distribution under the usual assumption - are standardised (subtract the expected value and divide by the standard deviation), we obtain standardised residuals, which are asymptotically normally distributed. They can therefore be subject to the usual interpretation from hypothesis testing and also to the usual critical values, e.g. 1.65 or 1.96 at 10% or 5% risk.

Adjusted residuals further correct for unequal margin dimensions and some researchers have shown that they are more appropriate than the usual standardised residuals, which is our recommendation, so we use adjusted residuals in our analysis (coloring of cells).

The 1KA application uses and colors the 1.0, 2.0 and 3.0 margins for the values of the adjusted residuals, which therefore roughly indicate the strength of the correlation in a given cell or the strength of the deviation from the null hypothesis assumption. Meaning of the values for the standardised residuals:

  • above 1.0 implies a certain increase and attention,
  • above 2.0 (a simplification of 1.96) implies a statistically significant difference (sign< 0.05), i.e. the residuals differ from zero with a relatively small risk.
  • above 3.0 already implies a strong deviation (sign<0.01), which means that the residuals are almost certainly different from zero and therefore something is "happening" in the cell.

Cells colored blue mean that there are fewer units in the cell than expected, and cells colored red mean that there are more units in the cell than expected.

For example, if there are 30 units in a cell and the expected value is 20, the basic residual is 10. If, for example, gender and agreement/opinion are considered, we therefore say that e.g. men are significantly more FOR than we would expect if gender had no effect. If we subtract the expected value from the residual 10 and divide by its square root (the square root of 20 is 4.5, since the Poisson distribution has an expected value equal to the variance), we get the standardised residual, which in this case is greater than 2, since we have (20-10)/4.5>2.0.

If we correct this slightly on the basis of the formulae in the appendices below, we obtain an adjusted residual which - barring really extreme asymmetries in the margins (YES:NO, male:female) - has a quite similar value. In any case, we can conclude that there are statistically significant deviations in this cell, and on this basis we can also proceed to a substantive interpretation (e.g. reasons why men are more FOR).

The coloring of cells in 1KA is indicative, simplified and intended purely as a screening (exploratory) analysis. In the formal interpretation, either the exact standardised or - better still - the adjusted residual is provided and interpreted in the usual sense as the examples below indicate.

The exact residuals are obtained in 1KA by selecting the checkbox option to calculate them (next to the independent and dependent variable dropdowns). 

Of course, the whole table and its Chi-square can be interpreted. But - as mentioned before - the residuals are more precise than the full Chi-square because they focus on exactly each individual cell where outliers occur. Further insight is gained by analysing the difference in shares based on a T-test.

Of course, all of this together is only valid for nominal variables. If one of the variables is "well" ordinally ordered - and even more so if there is a definite interval or ratio scale - we would, of course, prefer to use a T-test or analysis of variance.

Some useful links:

Related content

1KA is free to use for basic users