Tuesday, December 11, 2007
Chapter 7 Random Variables
http://news.yahoo.com/s/ap/20071213/ap_on_re_us/hiv_lawsuit
This test will be on Thursday, December 13, 2007.
You will need to create a chart to remind yourself about the formulas for this chapter until you have practiced enough to know them by heart. A sticky-note at pages 396 and 400 would also be helpful!
Mu = the POPULATION average. This is a parameter.
Sigma squared = the POPULATION variance. This is also a parameter.
Standard deviation is the square root of the variance.
X-bar is the sample average, the unbiased estimator of the population average. It is a statistic.
S-squared is the sample variance, the unbiased estimator of the population variance. It is also a statistic.
HW due Wednesday, December 12th: 7.13, 7.15, 7.17, 7.34, 7.42
HW due Tuesday, December 11th: 7.24, 7.28. Use the formulas and the examples in the text.
HW due Monday, December 10th : 7.2, 7.4, 7.7.
Monday, December 03, 2007
Chapter 6 Probability
Previous tests will be returned to students as soon as they are graded.
Prepare for the test. Work problems from each section. Read the chapter and section summaries. Write down what you are doing. Draw the Venn diagram or the tree diagram for complicated situations. Ask questions on the blog. Take the practice test.
Here are some answers to even HW problems:
Problem 6.10 (a) S= {all numbers between 0 and 24}
(b) = {any whole number up to and including 11,000}
(c) S = {0, 1, . . . 12}
(d) S= {any dollar and cents amount up to [insert your maximum guess here]}
(e) S = {any positive or negative number}
Problem 6.12 Four outcomes for two coins: {HH, HT, TH, TT}, eight for three coins: {HHH. HHT, HTH, HTT, THH, THT, TTH, TTT}, and sixteen for four coins (do that one yourself!).
Problem 6.16 (a) YYY -0000 through YYY-9999 = 10,000 numbers. (b) YYY-ZXX-XXXX, each X having 10 possible numbers, except the number can't start with a 0 or a 1 meaning that Z has only 8 possible values, so this means 8 * 10^6 LESS the restricted numbers (911-xxxx, 411-xxxx, etc.)
Problem 6.20 P(moves to another class) = 1 - P(stays) = 1 - .46 = .54.
Problem 6.24 P(wins large battle) = .6, P(wins three small battles) = P(wins individual small battle)^3 = .8^3 = .512. Choose the strategy with the larger probability of occurring.
Problem 6.40
Venn diagram has two circles representing getting job A and getting job B.
Both jobs: intersection of the two circles, the overlapped part, the biscuit.
First but not second: the part of circle A that is not within circle B
Second but not first: the part of circle B that is not within circle A
Neither: the part that is in the background, in NEITHER circle.
Problem 6.44
P(W) = 856/1626
P(W given prof degree) = 30/74
These are not the same. so gender and professional degree are not independent
Problem 6.56
P(y <> x) = 1/8,
P(y > x) = 1/2,
P(y <> x) = P(y <> x) /P(y > x) = 1/4.
Problem 6.48: P(W) * P(Manager given W) = P(Woman AND Manager)
One pattern that shows up a lot is Marginal * Conditional = Joint
If you divide both sides by Marginal you get
Conditional = Joint / Marginal.
IFF means IF AND ONLY IF.
IFF P(A) * P(B) = P(both A&B), A and B are independent.
IFF P(A) = P(A given B), A and B are independent.
HW for Tuesday night: DO problems 6.33 and 6.48. Read 6.66 and be prepared to work the problem. Essential question: How do mathematical independence and our regular understanding of independence relate? The chapter 4 tests were returned today. HW for Monday night: 6.39, .40, .53, and .56. The problem we worked today in class was problem .65. You would be wise to work through this problem and problem .66.
Notes from Friday (11/30) are embedded in the purple sections below.
Conditional probability rules:
PLEASE NOTE THE CORRECTION! BLOGGER WON"T ACCEPT THE VERTICAL LINE SYMBOL!!!
P(A GIVEN B) = P(A and B)/P(B)
P(B GIVEN A) = P(A and B)/P(A)
so of course
P(B) P(A GIVEN B) = P(A and B), the joint probability of A and B. It may be helpful to think of it like cancelling factors in the numerator and denominator of a fraction EXCEPT that the result is the JOINT probability. Be careful.
P(A) P(B GIVEN A) = P(A and B), again, the joint probability of A and B.
These relationships can be represented in two-way tables, Venn diagrams, and tree diagrams. The count within a cell of a two-way table divided by the marginal total is a conditional probability. Likewise, the joint probability for that cell divided by the marginal probability is also the conditional probability.
Tree diagrams can be useful when you are trying to work the problems backwards.
I don't think that I made this clear in class today:
P(A) = P(A and B) + P(A and not B) = P(A)P(B given A) + P(not A) P(B given not A).
HW for the weekend is problems 6.44 and 45.
If A and B are independent, then P(A) = P(A GIVEN B) and P(B) = P(B GIVEN A).
Interpretation: If A and B are independent, then whether or not B happened has no relationship with whether A happened.
Likewise, if A and B are independent, then whether or not A happened has no relationship with whether B happened
Today we used a Venn Diagram, a two-way table, and a tree diagram to represent the outcomes and probabilities associated with throwing two strangely-marked dice. All of the methods yielded the same answer.
HW for Tuesday (11/27) night: Re-work the weird dice problem from the AP exam (the one with two dice, one has only 9s and 0s, the other has 11s and 3s.) This time, instead of using simulation, use formal probability rules and a tree diagram, table, or Venn diagram. Answer the question in complete sentences. For part B, reconcile the answer with the joint probabilities you found in part A. Figure out the guidelines in your own words that tell you whether a price/reward is fair.
Get the reading done! What are the big concepts?
----------------------------------------------------------
Sorry for the delay: just got home from KSU.
HW for Monday night, Nov 26: 6.19, 20, 21
plus. . .finish State of Fear.
-----------------------------------------------------------
Add problems 6.24 and 6.25, due Monday, November 26.
Don't forget to read State of Fear.
------------------------------------------------------------
he complete list of HW problems due Tuesday: 6.9, 10, 12, 13, 16 (plus any others you feel like doing).
Don't forget to read State of Fear.
------------------------------------------------------------
Events that are mutually exclusive ARE NOT independent.
Addition principle: P(A or B) = P(A) + P(B) - P(A and B)
Multiplication principle: P(A) P(BA) = P(A and B), the joint probability.
When B and A are independent, P(BA) = P(B), A happening or not has not relationship to B happening, so P(A) P(B)=P(A and B). THAT IS ONLY WHEN THE EVENTS ARE INDEPENDENT.
Key vocabulary
parameter
sample space
event
probability
joint probability
independent
---------------------------------------------------------------------------
This won't be so bad. The test will be Thursday, Dec. 6.
What are YOU doing to maximize your understanding of the material?
- Are you creating an outline of the chapter?
- Have you developed a glossary for the vocabulary and formulas?
- Have you worked all of the homework problems when assigned?
- Do you read the sections that relate to the homework?
- Are you part of a study group?
- Do you ask questions?
- Have you worked problems from a study guide?
- Have you worked the online quiz (see the link on the right panel of this blog)?
- Do you look for the similarities and differences in the ways data are processed?
- Do you work with problems long enough to understand why the formulas work the way they do?
- Have you made connections between current concepts and prior knowledge?
- Have you gone online to review concepts that you have forgotten?
Just "going through the motions" does not lead to the success that you desire in Advanced Placement courses. Take control of your learning.
Be safe.
Wednesday, November 14, 2007
Chapter 5 -- Producing Data
The two sweet diagrams of experimental design are on page 272 (Completely randomized) and page 280 (block design).
Blocking is a form of control. When a large number of your experimental units share some pre-existing condition that may make their responses to the treatment vary tremendously WITHIN the treatment groups, you will have a hard time differentiating between the results of the treatment groups. You would really prefer to have the differences in results BETWEEN groups to be big enough so you can make a decision about your comparison. To reduce this vaiability, you may choose to BLOCK by the nuisance variable (the pre-existing condition). Then you RANDOMLY allocate the experimental units in each block to the different treatments. If there are two treatments, then each block is randomly broken into two treatment groups. You proceed by running the experiment on each of the blocks individually.
Thursday and Friday (11/8-11/9) Finish writing up the experiment described below, using all the concepts of section 5.2. ALSO, answer the free response (FR) problems from 2001 and 2003 handed out in class. You definitely NEED to read the section of the book. These are NOT opinion problems.
The Chapter 5 test is Thursday, November 15. Freakonomics must be read by 11/13.
Wednesday night's (11/7) HW: 5.65 PLUS design an experiment to answer this question (at least 6-7 sentences!!).
Does the choice of presentation technology make a difference in student achievement in a geometry class?
Conditions: Geometry classes at Lassiter
four teachers teach geometry
some teachers have students write HW on the board
some teachers have students write HW answers on the overhead projector
some teachers put their official answer transparencies on the O/H.
Two document cameras are available to use (Google document camera if you haven't seen one!)
Students are already assigned to the classes.
How could we design this experiment to answer the question? What questions or clarifications do you have? Bring at least seven complete sentences of helpful guidelines for performing this study.
Friday night's HW: 5.63 and 5.64. Complete most of your Freakonomics assignment this weekend. When you get to the part where the authors belabor their unique name theory, you can consider your assignment completed. what was your favorite part? What connections did the authors make that you agree with? that you don't agree with?
Thursday night's HW: 5.60 and 5.61
Wednesday night's HW: 5.54, 5.55, 5.56 Be safe.
Tuesday night's HW: Complete both of the problems from the 2001 exam.
Example of using the TORD to simulate a bag of M&Ms with the OLD color distribution:
Old Distribution:
Brown 30%
Red 20%
Yellow 20%
Green 10%
Blue 10%
Orange 10%
Let's try this two ways. First, let's use two-digit numbers to simulate candies according to the following schedule.
01-30 Brown
31-50 Red
51-70 Yellow
71-80 Green
81-90 Blue
91-00 Orange
There are no excluded numbers. If we draw the same number twice, use it again!
Using the following line from a table of random digits, simulate drawing 5 candies.
63996 32914
63>>>>Yellow
99>>>>Orange
63>>>>Yellow
29>>>>Brown
14>>>>Brown
The second way requires only one digit. Let 1-3 represent Brown, 4-5 for Red, 6-7 for Yellow, 8 for Green, 9 for Blue, and 0 for Orange.
21833 70905
Using the TORD above, you would get
2 Brown
1 Brown
8 Green
3 Brown
3 Brown
Link to interesting site about the Dewey-Truman polling error. Did you know who the third party candidate was who threw the wrench into the process? Strom Thurmond. Your parents will be impressed that you know this.
http://www.hannibal.net/stories/101998/Pollstersrecall.html
Interesting historical link about Tukey. Scroll to the middle to see his influence in predicting outcomes of elections.
http://www.amstat.org/about/statisticians/index.cfm?fuseaction=biosinfo&BioID=14
The two books I assigned for November are Freakonomics and State of Fear. Freakonomics discusses a lot of associations/correlations that promote critical thinking. State of Fear makes you enlightened consumers of research (even though it IS fiction). Many parents have probably already read one or both of these books. Last year's students (generally) loved them.
Wednesday, October 24
Take notes on the first section of the new chapter, especially new vocabulary.
For Monday and Tuesday of next week:
5.1-5.5, 5.8, 5.11, 5.17-5.18, 5.22, 5.23
Key concepts covered in class (alliteration, anyone??) today included
undercoverage
non-response bias
response bias
convenience sampling
voluntary response sample
and examples like the C-SPAN and American Idol calls, surveying the people sitting around you, the Dewey Defeats Truman mistake, answering with un-truths, failure to respond to surveys.
Can you match the concept to the example? Can you think of another example of each concept in action? Why does each of these result in data we cannot rely on?
Monday, October 08, 2007
Non-linear relationships
Assignment for Monday: Create an outline of the key points in the chapter, including all vocabulary words.
Also, do problems 4.54, 4.56, 4.62 (this study just celebrated its 20th anniversary!!!), and 4.66
So, how about those marginal and conditional distributions for two-way tables, huh?
The marginal distributions are the percents that each column or row represents in the entire table. For instance, if the total of one row was 250 and the total for the table was 1000, the marginal distribution for that row is 25%. You would continue to calculate percents for all of the other rows or the other columns--whatever the question asked for.
For the conditional distributions, you only consider a portion of your population, for instance only a specific row or column. Then, what portion of the observations recorded in tht small group shared the desired characteristic?
If there were 15 sophomores taking AP Chinese and 500 sophomores in a school of 2000 students, GIVEN THAT a student is a sophomore, the percent who are taking AP Chinese is 100*15/500.
Tuesday. October 16
Assignment for Friday: 4.34, .36, .38, .39, .40, .42, and .43
How did you like the Simpson's Paradox activity today?
When breaking data into two or more divisions by a lurking variable changes the "decision" for EVERY ONE of the sub-groups, the result is a Simpson's Paradox. For instance, the example today presented no clear, justifiable answer about whether we should fund Bolgg's Panacea or not.
The example in the book about the hospitals is instructive.
Good luck on the PSAT.
Monday, October 15
4.22-4.24, 4.27, and 4.28
Review the topics and procedures on the notes handed out today.
Have you ever heard of Simpson's Paradox???
Thursday, October 11
We've transformed non-linear data to a linear form, found the LSRL through the data, re-written the equation reflecting the nature of the lists used to develop the LSRL, and re-transformed the equation to model the original data.
You should have done problems 4.6 and 4.9. For tonight, DO problem 13 and READ ACTIVITY 4 and problem 4.15. If you feel excited about the investigation, read problem 16 also.
Monday, October 8
We graphed some relationships between x and y to determine whether we were allowed to run the LSRL on the data. Of course, we ONLY run the least squares regression on data that look like they have a linear pattern.
When the pattern in L1 and L2 looked like an exponential growth or decay model, we took the log of y in order to un-do the exponential. Putting the log y into L3, we proceeded to verify that the graph of L1 and L3 was approximately linear. We then ran the LSRL through that set of points.
The equation we found by using the LSRL will not run through our curve-y data, so we have to un-transform the equation. For the exponential case, we had used the log of y instead of y itself when finding the LSRL (but the original x values!), so we re-write the equation as log y-hat = a + bx.
We solve for y by taking the antilog of both sides ("ten-to-the" or 10^stuff). The resulting equation for y can be graphed with the original x and y data and should match the pattern pretty well.
If the model looks like a quadratic, square root, or other power function, you'll need to perform mostly the same functions, but on the logs of both x and y. The linear equation that passes through the straightened data will be transformed like this: log y-hat = a + b times log x, and the result will have a factor equal to x to the b power.
HW: Problem 4.1. The answer is in the back of the book, but there are a lot of sections to this problem.
Thursday, September 20, 2007
Chapter 3 Linear Relationships
Work AT LEAST 5 problems to prepare for Thursday's test.
Monday, October 1
3.50 and 3.52
Test on Thursday
Friday,September 28
Yo, this is David T.
Well today we broke up the deviation in a prediction, part into y-hat minus y-bar and an error part y -y-hat.
We also explained the significance of r^2 which equals the portion of the variation in y which could have been predicted using the regression relation.
Remember that if r^2 is close to zero, then the points on the graph are crazy and scattered. If r^2 is close to one, then the graph and points are predictable and are linear.
The HW is 3.46 and 3.49. (YES, this means YOU!) You = David V.
Thursday, September 27
KEY FORMULAS
b = r * Sy/Sx
a = y-bar minus b* x-bar
y-hat = a + b* x
residual = actual minus expected = y - y-hat
If residuals are small and scattered, then the linear model is a good model. If there is a distinct pattern (if you could predict what the residual would be for a particular x-value), then the linear model is not appropriate.
Be sure to WRITE what you see in the residuals ("The residuals are small and scattered, so a linear model is appropriate" or not) and what effect that observation has on your model.
Also be sure to write out the description of the y-hat equation in words: "The predicted value of [insert y variable here] is approximately [insert y-intercept here] plus [insert slope here] times [insert the x variable here]."
COMMON ERRORS:
Failure to use LinReg(a+bx) L1, L2, Y1
Failure to check that the observed y values are close to the predicted y values.
Failure to use the same x and y in your stat plot that you used in your linreg equation. (causes graphs to not show up!)
HW problem 3.39.
Wednesday, September 26
Problems 23 and 31 PLUS find the Least squares regression line for the Archaeopteryx data.
Tuesday, September 25
We re-worked the HW from last night and extended the concept by investigating what happens when you calculate the correlation coefficient for non-linear data (Anarchy! Riots! Dogs and cats living together!). Although you CAN calculate a correlation coefficient for non-linear data, the results tells you NOTHING.
Key points to remember:
- -1<= r <= 1. Always. No getting around it.
- r is dimensionless. If you change units or perform a linear operation on all of the values of x, or y, or both, your r will not change!!! In fact, what happens when you switch the order of the variables and calculate r for L2 and L1????
- r is affected by outliers. They increase the standard deviation, which causes the denominator to be smaller, which causes the r to be closer to 0.
- r only gives you information about linear relationships. If it isn't linear, then this linear modeling is inappropriate.
HW 3.13 and 3.19. All about the archaeopteryx.
See you tomorrow.
Monday, September 24
Do problem 3.18. This is just like what we did in class.
Friday, September 21
Problems 3.1-3.3 and 3.5.
Fifth and sixth periods: You did a great job with all the distractions today. Thanks for trying to stay on task.
Good job, Trojans! You make us proud.
CiCi's on Sunday? 2-4.
Be safe.
A new Chapter!!!
Thursday, September 20
Copy the formulas and definitions from Chapter 3 into your notes.
Friday, September 07, 2007
Chapter 2 Probability Distributions
Tuesday, 9/18
Work three of the problems that you set up last night and in class today. ALSO, do problems 2.41, .42, and .43 completely.
What more do you need in order to be successful with this? Make a list! Outline the important concepts from the chapter!
Why is the normal distribution such a big deal?
Monday, 9/17
Select 15 problems from Chapter 2. Split your paper in half (hotdog style). Write the details of the GIVENS on the left and the items that the problem asks you to find on the right side. You do not have to solve the problems. Watch carefully for those cases where there are multiple requirements.
Your test is Thursday.
Wednesday, 9/12
UPDATE: I have posted a WORD document with hints on the homework site: classhomework.com. You'll have to enter the password and then click on the file name.
Post to the blog and let me know when you get it--but don't ruin the fun for the other students!
-------------------------------------------------------------------------------------------
You KNOW that the area under the curve in a probability distribution is always one (That's why you get 1 when you use normalcdf(-infinity, +infinity)). Go back to the first parts of Chapter 2 and review the characteristics of a probability distribution.
Then, for HW, find the values of x which represent the Q1, median, and Q3 of a triangular probability distribution that starts at the origin and ends at (4, ???). Yes, you have to figure out what the value of ??? is so the are under the curve is 1.
You will use the formula for the area of a triangle: A = (1/2) base * height.
There are many different ways of attacking this problem. How many can you find???
Tuesday 9/11
HW 2.24, 2.25, 2.30
If you don't understand something, ask a question on the blog. Coming to class unprepared is not an option.
Monday 9/10
HW 2.15, 2.16, 2.22, 2.23
Friday 9/7
You developed equations today to standardize observed values.
Your formula for the z-score was (observed x minus the average)/(standard deviation).
You also used a formula to find the value of x that has a certain z-value:
x = average value + (z-score)(standard deviation)
but you probably noticed that the second formula is redundant.
You also learned about the Empirical Rule.
http://www.stat.tamu.edu/~west/applets/empiricalrule.html
This ONLY WORKS with normal distributions and it is only an approximation. We will learn more precise methods next week.
If you're interested in a neat relationship that works for other types of distributions, check this out: http://www.stat.tamu.edu/stat30x/notes/node33.html
Another neat website:
http://people.hofstra.edu/stefan_waner/Realworld/Summary7.html
Now, is anyone out there planning to CiCi's this Sunday? I won't go unless there is interest, so post your plans!
HW: 2.6, 2.7, 2.8 and read up to that point in Chapter 2. Be safe.
Wednesday, August 29, 2007
Graphs and standard deviation & CiCi's
Your test is Thursday. Begin to prepare now by working problems and creating an outline.
Homework: For those in class today - rewrite your responses to the FR questions, plus work problem 1.4 from the text.
For those absent today: Problem 1.4 PLUS ALL OF 1.48-1.52. Pick up your original responses to the FR upon your return to complete overnight. If you were participating in Senior Skip Day, your absence is unexcused.
CiCI's
YES!!!! There is a request for CiCi's this Sunday, so I will be there from 2 to 4. That is the one by the Walmart at Trickum and 92 (close to Arby's).
Test is Thursday, 9/6!
If you are going on the marketing fieldtrip, stay after school to take the test in room 214 at 3:30. Don't forget to bring your calculator.
For Tuesday (9/4), select one odd and one even problem from the set 1.48-1.52 and work them completely. Become an expert on one of the problems.For Friday (8/31), complete problems 1.41 and 1.43. The answers are in the back of the book, but that is not sufficient! You must show all work and explain your actions.
By Thursday (8/30), you should have both the graph from the Internet, a newspaper, or magazine and problems 1.35 and 1.36 from the text.
RE: the graph
You will identify the variable(s) represented in the graph and the type of graph you brought. Are the data numerical or categorical? Are numerical data discrete or continuous? Does your graph represent one variables or two? Is a trendline appropriate for your data?
RE: Standard deviation
The standard deviation of a sample of data is like an average deviation from the sample mean. It is the square root of the sample variance, which is an unbiased estimator of the population variance.
If we just found the sum of the deviations, we would get a sum of zero because some data are above the mean and some are below. Because of the definition of the mean, the positives and the negatives cancel each other out.
Instead, we square each deviation so the numbers we add together are all positive. We "average" these numbers by dividing by (n-1). You remember that n is the number of observations. We subtract one because we are using an estimate derived from the data themselves for x-bar. This gives us the sample variance or s-squared. To get the value of s just take the square root.
In formula form, s = sqrt(sum of all the squared deviations/(n-1)). The formula for the first squared deviation is (x minus x-bar)^2. Again, x-bar is the average of the x values.
The same relationship holds between sigma and sigma squared, the population standard deviation and the population variance: you take the square root of the variance to get the standard deviation.