Posts

Showing posts with the label statistics

Forcing an xy scattergraph for short columns in Google Sheets

Image
In a data exploration involving hoop rotation period versus diameter Google Sheets does not always produce the desired scatter graph. The issue appears to small sample size or short columns.  Spreadsheet Selecting a scatter chart for this data results in the following chart for some students:   Note the monotonic spacing on the x-axis. Google Sheets has defaulted to a form of line graph with monotonic spacing on the horizontal axis. To reunite the paired data, select Use column A as label under edit chart on a laptop or desktop computer.  Note that this option is not accessible from the Google Sheets mobile app. Of note is that while working inside the Google Sheet chart module in the mobile app the chart appears to configure as expected. Once one exits the chart module, the columns split into two data sets. There is no known easy fix on mobile. The repair essentially requires getting onto a laptop or desktop.  There are hackarounds. One is to manually type the data ...

DDFT Statistical foundations for problem based learning

Image
The design intent outlined fall 2025 was a set of five case studies that the DDFT program students would engage with using problem-based-learning approaches. The original plan was to offer these sessions spring 2026 but scheduling issues prevented the sessions from occurring. The sessions were moved to summer 2026 with the intent that this would be a residential, problem-based-learning set of sessions. The first day of the sessions saw an attendance of one student and notes from many others that they were either off-island, or had conflicting class schedules, or simply intended to engage with the material online Elizabeth Synchronous video conferencing sessions are notorious for attendance issues, and have all of the same time conflict issues as a residential section. Add in connectivity and power issues, and synchronous sessions are generally a nonstarter. How to engage in asynchronous group based PBL online is not clear. Stack on top of this that the medical case studies ...

12.3 FSM Family Health and Safety Study data exploration

Image
In chapter 12.3 the textbook recommends considering the statistical tools that you already know how to use. You can be 95% confident that the instructor has chosen a problem that can be resolved by the tools taught in the course. In the "wild" there are many more tools to consider. F-tests for a difference of variances (standard deviations), confidence intervals for a slope, tests for differences of medians, tests for normality, or Bayesian credible interval. All of these are beyond the scope of this introductory course. Thus the student is left with basic statistics (chapters one, two, three), correlations (chapter four), confidence intervals (chapter nine), hypothesis tests against a known mean (chapter ten), and tests for a difference in two sample means (chapter eleven). Those are the tools that have been covered, in the course an open data exploration exercise is limited to using those same tools.  The course features reduced content to provide more time for data explor...

Google Sheets one function to rule them all and in the darkness bind them

Image
I first began using the spreadsheet software Lotus 1-2-3 in 1986 and have memorized and mastered an ever increasing number of spreadsheet functions over the decades since. Google has introduced a new AI function in Google Sheets that acts like the one ring in the Lord of the Rings: one function that rules all other functions.  I spend a couple of weeks in MS 150 Statistics covering the normal distribution, standard error of the mean, t-distribution and t-critical, and 95% confidence intervals. The students learn to calculate the mean, the standard deviation, the standard error of the mean, t-critical, and then the lower and upper bounds of the 95% confidence interval. The Google AI function replaces all spreadsheet functions with a single function () and a prompt inside the function. Tell the function what you want, and the function constructs the functions that give you the answer. I have only entered a prompt inside the new AI function in cell C1 with data in column A of the...

Federate States of Micronesia 2023 census released

Image
I was fortunate to get to attend the release of the 2010 FSM census sixteen years ago. I was fortunate to obtain and retain the data released that year, data which subsequently became unavailable.  This year a slide deck was prepared to share the results with the statistics course students.  Having the values available from 2010 along with data from 2023 provides some insight into where the population decreases have been the largest. Outmigration is under these population losses. The four states.

Sending small datasets in a URL using ZipTable

Image
MS 150 Statistics often uses small datasets that are usually shared with students via copy-on-link of Google Sheets. ZipTable provides another way to move small datasets that is ideal for communicating datasets in limited bandwidth environments. For example, a small dataset extracted from the 2023 Federated States of Micronesia can be encoded and shared as a link . After opening the   link  in ZipTable, the student can download the data as a CSV file. That file can then be uploaded into Statisty.app for analysis. This allows everything to happen in a browser - no other software required. Neither ZipTable nor Statisty require logging in. 

11.1 Paired marbles t-test

Image
The marbles were pre-split on 5.1 grams and 5.2 grams. Less massive on one side, more massive on the other.  This term the students detected a mass difference with the results being statistically significant. The students are always surprised when they get the mass difference correct. One only has a sensation of a faint difference, but no surety.  Yet 14 of 18 were correct with one tie and three incorrect, far from the 9|9 split that might be randomly expected.

10.2 From confidence intervals to hypothesis testing

Image
Section 10.1 introduced using confidence intervals for hypothesis testing.  Spreadsheet In the above fibobelly exercise the tested value was 1.618, the golden ratio. There exists a theory that that the distance from the ground to the belly button divided by the distance from the belly button to the top of the head is the Fibonacci ratio 1.618.  The female confidence interval did not capture the test value of 1.618, the male confidence could not rule out the golden ratio as being the male fibobelly ratio. Desmos The shift to 10.2 hypothesis testing with the t-statistic was supported by Desmos files. Above the t-statistic for the females is out beyond the upper t-critical. Values from the spreadsheet were entered into Desmos. Desmos For the males the t-statistic is not beyond the t-statistic.

9.2 Paper aircraft flight distances

Image
On the way to work the idea of throwing the paper aircraft off of the mezzanine crystalized when the resolution to the lack of a pre-existing population mean was realized to be having the students predict the flight distances.  The mezzanine is very nicely appointed at this time. The students were asked to guess distances in feet as this would be the measurement system with which they are most familiar. To help the students visualize the space, the above image was prepared. Note that lanai is a Hawaiian word. The student union has roughly the layout of a nahs, a U shaped two level structure. The equivalent place in a nahs is called either nan kadei (towards the high titles at the front) or nan pahpei (towards the back, towards the apwin dies). Note that these are 1978 old system spellings. Reference the Pohnpei cultural documents specifically page five and eight .  The tape measure is anchored by the Tripltek directly under the mezzanine r...

8.2 Five marbles and the standard error of the mean

Image
Materials required are marbles and a 200 gram scale. This term a mix of marbles was used.  Students were told to pick five marbles. This avoided an issue that had been created by giving students five of the same marbles. When different sets of five were in use, some were the smaller, lighter blue marbles and these would generally not capture the population mean. A spreadsheet , single tab for the Smartboard memory limitations, is used to generate data. Tallying a dozen students took until around 9:25, this photo was taken at 9:32.  After the first five marbles were massed the point estimate of the population mean was 5.06. The massing was paused to explain that this was the best estimate of the population mean for that student. The point estimate for the population mean.  Masses are done marble by marble. After the second student the split in the point estimate was pointed out. And that both were now "wrong" - the population...

7.1 The shape of random variation

Image
The starting concept was to roll one, two, and then three dice mapping the results onto the white board to show the distributions. This morphed minutes before class into passing out a single die to 12 student pairs and having the students tally data into a shared spreadsheet.  The concept worked, with the complication that different individuals, and pairs, moved at vastly different paces through the data. With some working as pairs and others as individuals, there were only nine dice being rolled. Because the spreadsheet included the distribution for 12 dice, the data was padded using three columns containing the function =int(rand()*6+1). In retrospect this felt like a kludge.  Karen and Kimilane started as a pair, and then were each given their own column to add another student generated data column. The Smartboard proved problematic with high latency - adding a column was difficult. Perhaps this was in part that nine groups were entering data to...