Assessing Learning in Introductory Statistics
MS 150 Statistics is an introductory statistics course with a focus on statistical operations and methods. The course is guided by the 2007 Guidelines for Assessment and Instruction in Statistics Education (GAISE), the spring 2016 draft GAISE update, and the ongoing effort at the college to incorporate authentic assessment in courses. A history of the evolution open data open data exploration exercises and associated presentations as authentic assessment in the course was covered in a May 2017 report.
Three course level student learning outcomes currently guide MS 150 Introduction to Statistics:
The use of prior final examinations as a practice tests was made possible in part due to the use of Google Sheets for data sharing starting in spring 2017 and the continued use of Google Sheets in subsequent terms. Spring 2017 was the first term to use Google Sheets as the supporting software for the course from day one.
Spring 2018 saw the adoption of Schoology Institutional and the further deepening of integration between the course and Google Sheets via the Google Drive Assignments application in Schoology.
The basic structure of the final examination has been fairly stable over time, and performance on an item-by-item basis is also fairly stable over time.
Performance on the final examination saw an adjustment back towards the longer term average, a statistical effect known as returning to the mean. A linearly weighted running average suggests that the mean to which the final examination is returned has been rising since 2012.
The course average since 2007 has also remained stable and has tended to remain within seven percent of 79%. This term's 84.7% course average remains within the historic range. Continuous improvement is a laudable goal, the reality is that values return to long term means. The uptick seen this term is likely attributable to the strong performance of the DDFT students in the historically weak 8:00 section. The course average is predicted to drop spring 2020.
A student presents atmospheric CO₂ versus temperature on Pohnpei
- Perform basic statistical calculations for a single variable up to and including graphical analysis, confidence intervals, hypothesis testing against an expected value, and testing two samples for a difference of means.
- Perform basic statistical calculations for paired correlated variables.
- Engage in data exploration and analysis using appropriate statistical techniques including numeric calculations, graphical approaches, and tests.
The course wrapped up coverage of content five weeks prior to the end of the term. This was a week earlier than the previous term, but one of the weeks in the final five was the effectively lost to a set of early November holidays. The content was compressed by one day over prior terms by the dropping of the material in section 9.12 in the textbook. The material in 9.12 repeatedly led to students making errors in subsequent calculations of confidence intervals for sample sizes less than 30. The content was also compressed a further day by combining sections 9.2 and 10.1.
The calculation of the margin of error for the mean was also dropped from the curriculum as the calculation added no further insight into confidence interval calculations over using t-critical multiplied by the standard error. The margin of error was not adding anything to the course, and the linkage of two standard errors to the margin of error is more an artifact of historic simplification of calculations of a 95% confidence interval. The students would glom onto the two and even after learning about t-critical would return to using two in subsequent calculations. This term the option of using two for sample sizes above thirty was never mentioned. This then disconnects confidence intervals from ordinary and extraordinary z-scores, but that connection was tenuous at best.
Final examination details and performance
In the final five weeks, the students engaged in a series of four data analysis, exploration, and presentation exercises. The course then ended with a final examination which was completed as an online test inside Schoology.
Practice tests from prior terms were posted in a time staggered order beginning two weeks prior to the final examination. The practice tests were not modified from their original content, thus the margin of error calculation remained in these practice tests although the topic was dropped from the course this term.
The use of prior final examinations as a practice tests was made possible in part due to the use of Google Sheets for data sharing starting in spring 2017 and the continued use of Google Sheets in subsequent terms. Spring 2017 was the first term to use Google Sheets as the supporting software for the course from day one.
Spring 2018 saw the adoption of Schoology Institutional and the further deepening of integration between the course and Google Sheets via the Google Drive Assignments application in Schoology.
Forty-one students sat the final examination fall 2019. The final included 28 questions.
The table depicts the percent success rate on each item on the final examination across nine terms.
Students evidenced strength in calculating basic statistics of both one and two variables. Performance on paired variable statistics was on par with prior terms, with the exception of calculation of the correlation coefficient r which saw 100% correct answers. This occurred in part because the spreadsheet function for slope and intercept are column order dependent, the correlation is not dependent on the order in which the columns are entered.- Part I: What is the sample size n?
- Calculate the mode:
- Calculate the median:
- Calculate the mean.
- Calculate the minimum:
- Calculate the maximum:
- Calculate the range.
- Calculate the first quartile Q1:
- Calculate the third quartile Q3:
- Is this the correct boxplot?
- Is this the correct histogram?
- Calculate the sample standard deviation sx.
- Calculate the standard error SE of the sample mean.
- Calculate t-critical for a 95% confidence level.
- Calculate the lower bound for the 95% confidence interval for the population mean μ.
- Calculate the upper bound for the 95% confidence interval for the population mean μ.
- Part II: What is the sample size for this paired data?
- Calculate the slope.
- Calculate the y-intercept.
- Calculate the correlation coefficient r.
- Is the correlation none, weak/low, moderate, strong/high, or perfect?
- Based on the correlation, does my cadence appear to relate to my speed?
- Part III: Calculate the p-value for a difference in the pairwise mean between the two samples.
- Is the difference in the means statistically significant at a risk of a type I false positive error of 5%?
- Based on the analysis and a risk of a type I error (alpha) of 5%, do you reject or fail to reject a null hypothesis of no difference in the means?
- Calculate the pooled standard deviation.
- Calculate the effect size.
- Is the effect size small, medium, or large?
The table depicts the percent success rate on each item on the final examination across nine terms.
Databars for the data
The MS 150 Statistics course fall 2019 consisted of two sections, a total of 44 students enrolled as of term end, 23 females and 21 males. As noted above, of the 44 students, 41 sat the final examination. One student had left the island, two other students had stopped attending the class after midterm. The two sections were kept in curricular synchronization during the term. Both sections covered the same material, worked the same assignments, and gave presentations on the same topics. The sections met at 8:00 and 9:00 on a Monday-Wednesday-Friday schedule.
There were 23 females and 21 males in the two sections of statistics. Gender differences were not statistically significant for the course nor for performance on the final examination.
1.0 Perform basic statistical calculations for a single variable up to and including graphical analysis, confidence intervals, hypothesis testing against an expected value, and testing two samples for a difference of means.
2.0 Perform basic statistical calculations for paired correlated variables.
3.3 Draw conclusions based on statistical analyses and tests, obtain answers to questions about the data, supported by appropriate statistics
Student learning outcome performance on basic statistics (course learning outcome one) has been stable at around an 80% success rate. The success rate of 80.5% for student learning outcome one this term was echoed in the 80.8% success rate on items testing basic statistics on the final examination. Statistics is traditionally considered a difficult and challenging course, but as one student responded, During the first day I remember thinking, "Only 'I will fail this class. It looks so hard.' But no, if we focus and listen we will be okay."
Student learning outcome performance on paired statistics and linear regressions (course learning outcome two) has been stable at around a 78% success rate. The success rate of 78.5% this term for student learning outcome two this term was also seen in performance on the final examination where there was a 75.6% success rate on paired statistics items.
Student learning outcome performance on data exploration, analysis, and use of appropriate statistics (course learning outcome three) has been evaluated by in-class presentations given by the students. Course learning outcome three has been less stable with a multiterm average of an 81% success rate. This term the success rate slid to 69%, down 4% term-on-term from 73% the prior term. This value, however, may not be comparable across the seven terms. In the fall of 2016 the course had only four data exploration presentations, spring 2019 there were seven presentations, fall 2019 there were eight presentations. This increase in the number and complexity of data explorations is being driven both the GAISE recommendations from the ASA and by students reactions such as the following:
The activity that contributed the most to my learning was...
Performance by section on the final and in the course
This term the 8:00 and 9:00 section showed no significant difference in performance on the final examination. There was also no significant difference in the overall course average between the two sections. In prior terms the 8:00 course average is often statistically significantly lower than the 9:00 course average. This is usually attributed to transportation and attendance issues that are more pronounced in the 8:00 section than the 9:00 section. This term the 8:00 section was a section reserved to an on campus program, Doctors and Dentists For Tomorrow. The DDFT program students are pre-selected pre-medical students.
Performance differences by gender in the course and on the final exam
As taught, the course makes statistics accessible and provides opportunities for students to experience success in a mathematics course. As one student wrote on a term end reaction assessment, During the first day, I remember thinking…"that this course was going to be the hardest and that I will fail and not understand anything. Turns out that this was my favorite course and it kind of boost up my self-esteem when presenting."
Student learning outcome performance over multiple terms
The introduction of Schoology Institutional in January 2018 has made possible tracking of performance based on student learning outcomes. Prior to January 2018 Schoology Basic permitted the entering of student learning outcomes, but the Basic version does not provide access to the Mastery screen. Once the college adopted the institutional version, however, Mastery data from as far back as the instructor measured against student learning outcomes becomes available. Data across seven terms is reported for the following three learning outcomes:1.0 Perform basic statistical calculations for a single variable up to and including graphical analysis, confidence intervals, hypothesis testing against an expected value, and testing two samples for a difference of means.
2.0 Perform basic statistical calculations for paired correlated variables.
3.3 Draw conclusions based on statistical analyses and tests, obtain answers to questions about the data, supported by appropriate statistics
Performance on student learning outcomes over multiple terms.
Student learning outcome performance on paired statistics and linear regressions (course learning outcome two) has been stable at around a 78% success rate. The success rate of 78.5% this term for student learning outcome two this term was also seen in performance on the final examination where there was a 75.6% success rate on paired statistics items.
Student learning outcome performance on data exploration, analysis, and use of appropriate statistics (course learning outcome three) has been evaluated by in-class presentations given by the students. Course learning outcome three has been less stable with a multiterm average of an 81% success rate. This term the success rate slid to 69%, down 4% term-on-term from 73% the prior term. This value, however, may not be comparable across the seven terms. In the fall of 2016 the course had only four data exploration presentations, spring 2019 there were seven presentations, fall 2019 there were eight presentations. This increase in the number and complexity of data explorations is being driven both the GAISE recommendations from the ASA and by students reactions such as the following:
The activity that contributed the most to my learning was...
"Doing open data explorations and presenting them, because it was something I had to do on my own to test my knowledge on what I've been learning in this class and presenting has been something fun for me."Over the past seven terms the open data explorations have increased in number and in the level of statistical difficulty that the data presented to the students.
"Presentations because I learned how to solve the basic statistics on my own and to present well."
"Presentations. They helped me apply my knowledge and skills from what I have learned and heard from the instructor. It make me use my brain. It was quite difficult and I always made mistakes but I learned through them. It was a good challenge, I should say."
Longer term trends
As an educator, I am aware that there is a penchant in education for "continuous improvement." The reality is that there is far more inertia in a value than I suspect those in education comprehend. Success rates tend to be stable over long periods of time and reflect both the difficulty of the material as well as the many reasons students do not succeed on the material. For every student who did not do well, there is a complex back story. Data for success on the final examination demonstrates this longer term stability and the tendency to return to the long term average.
Long term lack of a trend in final examination averages Fall 2005 - Fall 2019.
Note that the y-axis does not start at zero nor end at 100, the vertical range is exaggerated.
Performance on the final examination saw an adjustment back towards the longer term average, a statistical effect known as returning to the mean. A linearly weighted running average suggests that the mean to which the final examination is returned has been rising since 2012.
Long term course average and standard deviation
The course average since 2007 has also remained stable and has tended to remain within seven percent of 79%. This term's 84.7% course average remains within the historic range. Continuous improvement is a laudable goal, the reality is that values return to long term means. The uptick seen this term is likely attributable to the strong performance of the DDFT students in the historically weak 8:00 section. The course average is predicted to drop spring 2020.
The standard deviation of the course averages is also relatively stable remaining within four percent either way from the long term average of 15%. Fall 2019's 12% continues this stability in variation.
Overall, given a list of numbers and spreadsheet software, students show a strong mastery of basic statistics, good capabilities with linear regressions, and more moderate abilities with confidence intervals and open data exploration.
Recommendations for the spring 2020 term include retention of the increased number of presentations and continuing the combining of sections 9.2 and 10.1. The textbook could be updated to reflect the dropping of section 9.12, section 10.1 should probably be moved into chapter nine after section 9.2. Retention of the Wednesday to Monday presentation cycle should also be considered for spring 2020.
Overall, given a list of numbers and spreadsheet software, students show a strong mastery of basic statistics, good capabilities with linear regressions, and more moderate abilities with confidence intervals and open data exploration.
Recommendations for the spring 2020 term include retention of the increased number of presentations and continuing the combining of sections 9.2 and 10.1. The textbook could be updated to reflect the dropping of section 9.12, section 10.1 should probably be moved into chapter nine after section 9.2. Retention of the Wednesday to Monday presentation cycle should also be considered for spring 2020.










Comments
Post a Comment