Posts

Showing posts with the label statistics

5.3 50 orange 35 blue 15 white marbles

Image
One hundred marbles. Students blind draw a single marble. Population probability is known. Tally the sample to obtain relative frequency.  Last spring the recommendation was to move to ratio level data . Upon further thought, the ratio level data would not demonstrate the linkage between mathematically determined probabilities and statistically inferred probabilities. Ratio level data wouldn't have obvious population values.  One could use the mass of the marbles to predict a known population mean, but that is a conceptually harder system in which to explain what the probability of individual mass classes ought to be. Those weight class probabilities would be driven by an underlying normal distribution. Students would not comprehend why the predicted probability for the class bounded by the mean minus a standard deviation and the mean has a predicted probability of 34% Fifty orange marbles, thirty-five blue marbles, and fifteen white marbles sets up an easier to se...

4.2 4.3 Slope intercept and slope

Image
Due to 4.2 and 4.3 being combined as the result of a holiday, swizzle data was gathered on the porch to reduce travel time. The period was still tight. A ten meter run on the porch was still functional. Spreadsheet Data was recorded directly into the spreadsheet on the Samsung Active5 tablet while on the porch. Another spreadsheet  tracks 163 runs back to 2011. The board was used to emphasize that the columns were reversed.  Note that the board is using the data from 4.1 in the previous class . This run would be particularly useful because of the large gaps in the split data out beyond 33 meters. That data can be extrapolated beyond maximum x.  The swizzle data is more useful as extrapolation below minimum x or above maximum x are physically unrealistic. Packing the demonstration, 4.2, and 4.3 was tight. Too tight to demonstrate the 4.1 homework and hooping data exploration. 

4.1 4.2 RipStik run

Image
The RipStik run returned to the covered walkway after testing out the new health center fall 2025 and the unfinished student center spring 2026. An 81 meter run from LRC to the A building was laid out. This was a cold open on the sidewalk with a note to students ahead of class to meet on the sidewalk. This saved significant time. Initial roll was called on the sidewalk. O meters was at the LRC.  Every three Meyers was marked out to 33 meters in front of faculty office. 51 meters at the junction. 71 at the turn 81 at the A building. Spreadsheet and Desmos  were prepared in advance. The layout worked well to illustrate the linearity present in the first 33 meters. The gaps to 51m, 71, and 81 meters provided a chance for interpolation. This also was very functional when riding the RipStik. the segment into 51 meters allowed for setting up for the slope into 71 meters. Speed increased to 2.25 meters per secon...

3.4 Shape of distributions

Image
An air conditioner iced up over the weekend and flooded the floor. The in floor electrical outlets presented a potential shock hazard, so MS 150 Statistics moved to A101. A symmetric heap was the predicted shape. Kozel measures Mary Bella Tape measures were used, with one inch under the foot, and then an inch added to measurement. Maireen working on getting measured Data was then recorded on the board by the students. Data on the board. The data on the board was entered into a spreadsheet . A lecture on the possible shapes followed. For the uniform shape a single die was simulated with 100 rolls. This was extended to six die and 500 rolls, the sum of which was nicely normal once the buckets were set to integer values rather than fractional boundaries. 

3.1 into 3.3 Charts

Image
Monday began with an unshared 3.1 Smartboard fall 2026 spreadsheet. This probably proved confusing and was probably too long. Left insufficient time to demonstrate how to Pareto sort data on the tablet as shared to the Smartboard. Simplification to just census data charts might be more appropriate. Cover when to use a circle chart and when to use a column chart. Wednesday was blue Wednesday. Although black smashed blue two-to-one in a 6-3 blowout. Attire was blue on top, black pants, green t-shirt, red shoes (brick red technically), and brown socks. Which captured the top five, although not only did black beat blue, but brown bested red as well. Still, the top five were arguably well met. Sample size and relative frequency were both brought in by this example.  Wednesday demonstrated nominal level histograms in Google Sheets and wrapped with the multi-term spreadsheet .  Friday ratio level histograms used the marble masses data . The 15 marble ran from 4.6 grams t...

Using Gemini notebook to check answers in statistics class

Image
As part of an introduction to MS 150 Statistics I showed the students how to enter a few basic statistical functions into a spreadsheet . I then used a Gemini notebook (formerly NotebookLM) preloaded with the textbook chapters to check the work. I started with a prompt: I want to give you some data followed by some statistical questions along with my answers. I want you to check my answers. The data: Car speed in mph 21.2 21.2 . . . 21.2 20.0 The questions: Statistic to calculate Sample size n Mode Median Sample mean x̅ Answers 18 21.2 20 20.3 Note that in the above the data was clipped, the original prompt included all of the data. Gemini/NotebookLM responded with the following analysis: Gemini notebook correctly discerned that the data set was multi-modal. There is a mode.mult function that can be used to return values in a polymodal data set, but that is not usually covered in the course.  The takeaway is that Gemini notebook correctly parsed and responded to a multipart prompt...

Mars and Murrie discovery learning in statistics

Image
This term a course introduction presentation was used to organize the material on the first day. The session included a demonstration of entering formulas into a spreadsheet . A demonstration of NotebookLM/Gemini notebook was done after the spreadsheet demonstration. In a notebook that was already pre-populated with the textbook chapters the following prompt was entered using shift-enter to prevent premature submission: I want to give you some data followed by some statistical questions along with my answers. I want you to check my answers. The data: Car speed in mph 21.2 21.2 . . . 21.2 20.0 The questions: Statistic to calculate Sample size n Mode Median Sample mean x̅ Answers 18 21.2 20 20.3 Notebook responded with the following analysis: After this demonstration the first data exploration was shown, MMs distributed, and massed using the scale. 

Forcing an xy scattergraph for short columns in Google Sheets

Image
In a data exploration involving hoop rotation period versus diameter Google Sheets does not always produce the desired scatter graph. The issue appears to small sample size or short columns.  Spreadsheet Selecting a scatter chart for this data results in the following chart for some students:   Note the monotonic spacing on the x-axis. Google Sheets has defaulted to a form of line graph with monotonic spacing on the horizontal axis. To reunite the paired data, select Use column A as label under edit chart on a laptop or desktop computer.  Note that this option is not accessible from the Google Sheets mobile app. Of note is that while working inside the Google Sheet chart module in the mobile app the chart appears to configure as expected. Once one exits the chart module, the columns split into two data sets. There is no known easy fix on mobile. The repair essentially requires getting onto a laptop or desktop.  There are hackarounds. One is to manually type the data ...

DDFT Statistical foundations for problem based learning

Image
The design intent outlined fall 2025 was a set of five case studies that the DDFT program students would engage with using problem-based-learning approaches. The original plan was to offer these sessions spring 2026 but scheduling issues prevented the sessions from occurring. The sessions were moved to summer 2026 with the intent that this would be a residential, problem-based-learning set of sessions. The first day of the sessions saw an attendance of one student and notes from many others that they were either off-island, or had conflicting class schedules, or simply intended to engage with the material online Elizabeth Synchronous video conferencing sessions are notorious for attendance issues, and have all of the same time conflict issues as a residential section. Add in connectivity and power issues, and synchronous sessions are generally a nonstarter. How to engage in asynchronous group based PBL online is not clear. Stack on top of this that the medical case studies ...

12.3 FSM Family Health and Safety Study data exploration

Image
In chapter 12.3 the textbook recommends considering the statistical tools that you already know how to use. You can be 95% confident that the instructor has chosen a problem that can be resolved by the tools taught in the course. In the "wild" there are many more tools to consider. F-tests for a difference of variances (standard deviations), confidence intervals for a slope, tests for differences of medians, tests for normality, or Bayesian credible interval. All of these are beyond the scope of this introductory course. Thus the student is left with basic statistics (chapters one, two, three), correlations (chapter four), confidence intervals (chapter nine), hypothesis tests against a known mean (chapter ten), and tests for a difference in two sample means (chapter eleven). Those are the tools that have been covered, in the course an open data exploration exercise is limited to using those same tools.  The course features reduced content to provide more time for data explor...

Google Sheets one function to rule them all and in the darkness bind them

Image
I first began using the spreadsheet software Lotus 1-2-3 in 1986 and have memorized and mastered an ever increasing number of spreadsheet functions over the decades since. Google has introduced a new AI function in Google Sheets that acts like the one ring in the Lord of the Rings: one function that rules all other functions.  I spend a couple of weeks in MS 150 Statistics covering the normal distribution, standard error of the mean, t-distribution and t-critical, and 95% confidence intervals. The students learn to calculate the mean, the standard deviation, the standard error of the mean, t-critical, and then the lower and upper bounds of the 95% confidence interval. The Google AI function replaces all spreadsheet functions with a single function () and a prompt inside the function. Tell the function what you want, and the function constructs the functions that give you the answer. I have only entered a prompt inside the new AI function in cell C1 with data in column A of the...

Federate States of Micronesia 2023 census released

Image
I was fortunate to get to attend the release of the 2010 FSM census sixteen years ago. I was fortunate to obtain and retain the data released that year, data which subsequently became unavailable.  This year a slide deck was prepared to share the results with the statistics course students.  Having the values available from 2010 along with data from 2023 provides some insight into where the population decreases have been the largest. Outmigration is under these population losses. The four states.

Sending small datasets in a URL using ZipTable

Image
MS 150 Statistics often uses small datasets that are usually shared with students via copy-on-link of Google Sheets. ZipTable provides another way to move small datasets that is ideal for communicating datasets in limited bandwidth environments. For example, a small dataset extracted from the 2023 Federated States of Micronesia can be encoded and shared as a link . After opening the   link  in ZipTable, the student can download the data as a CSV file. That file can then be uploaded into Statisty.app for analysis. This allows everything to happen in a browser - no other software required. Neither ZipTable nor Statisty require logging in.