• Bhangra discussion is still going strong. Join us in our Facebook group!

    New user registration has been closed (as it was entirely spam). We encourage you to post in our Facebook group, even if it's a followup to an existing thread. BTF will continue to be archived and hosted here - Saleem

Elite 8 Bhangra Invitational - Judging Package

Vick

Member
Messages
990
You guys are missing the whole point here... They got the winner right regardless of who was judging and how they judged... It's all about getting the right results... Yea the rotating panel of judges has issues, but if Nachdi had placed over VCU, this would have been a tragedy... They got it right, doesn't matter how you break it down...
 

kinnell

*Account Deactivated*
Messages
2,159
Vick said:
You guys are missing the whole point here... They got the winner right regardless of who was judging and how they judged... It's all about getting the right results... Yea the rotating panel of judges has issues, but if Nachdi had placed over VCU, this would have been a tragedy... They got it right, doesn't matter how you break it down...
A airplane tries out a new engine but it has a malfunction and crashes. Luckily no one is killed. It's all about the fact that no one died. Let's not look into the malfunction. Let's not see how it could be prevented next time. Let's not look into how it can be improved for next time. Don't get me wrong, the end result is very important, but Vick, let us nerds have some fun, ok? :)
 

Kapil

Wear Responsibly 2.0 || One for All...All for FAUJ
Messages
133
kinnell said:
Data: http://spreadsheets.google.com/pub?key=tmUnF56SYz-5N56K01QvDOQ&output=html

The inherent flaw in the system are the team judges. In my opinion, there are three reasons why:
1) They do not watch all the performances.
2) Their qualifications to judge are never questioned.
3) While we do want to overlook it, there are rivalries between teams. In addition, teams may not like each other's styles. Essentially, we have a variation of game theory present because these teams are competing against each other.

Let's compare how the teams ranked (both raw and cut) to how the actual judges (both individual and combined).

Attached.

It will be interesting to see how things would have played out if we normalize the scores.

EDIT: I made a mistake in the Totals. I fixed it now.
Some notes:

1) You can't say with 90% confidence you "know", you can only say that if you repeated this experiment with the same conditions 90% of the confidence intervals would contain the true mean. Someone else said that's not significant but it IS at an alpha of 10% (most common used is 5%, so in that case you're right).

2) Kinnel you are assuming that Bhangra judging as a whole is normally distributed, which it may not be. All judging may be skewed to the higher end (note no scores below 50 at all). Also, you can't use the Law of Large numbers or the Central limit theorem because your sample size isn't big enough for this example; you'd need many competitions' judging data. T-distr could be better since it's got fatter tails, but still as DF increases you approach the normal assumption.

3) If you were to mix game theory with statistical inference, but you would have to take into account probability of certain events specific to certain teams, thereby creating a more complex bayesian discrete distribution per judge on a team by team basis. Then you could compare distributions by assuming a hypothetical distribution, doing a siimulation of how he/she would judge at other comps, and compare the distributions vs other judges with a Chi Square Test due to large sample probability laws... I think....hah

Short version: Crapton of intricacies --> hard to actually model. You'd have to infer a judge by judge distribution, simulate, and compare if they are actually similar. One fact you missed out is that a judge may inadvertently make mistakes in #s due to subjectivity of writing down #s on the fly during judging, thereby giving us the "garbage-in, garbage-out" problem with the data; judges can't go back after they see other teams and "fix" the score.
 

yraparla

SwizzeeMusic.com
Messages
2,072
great points kapil, garbage in garbage out problem is exactly what a judge needs to judge all the teams, so that the relativeness of their scores can be complete

that's why i'm saying a relativistic model would be good here. Have the judges do their best to actually judge each team individually and allow the math to do the adjustments
 

Ekta Kaur

New Member
Messages
4
Meistro said:
However, I would make the grading scale to be more objective because there was clearly a discrepancy in some of the scores. Either some of you were being assholes on purpose, didn't care, or wanted to grade easy. There has to be a balance and all judges have to judge the same.
I completely agree, and regardless of whether this judging style "worked" or not, I didn't like it.
 

nb0913

New Member
Messages
280
trying to solve this problem using math isn't going to get us anywhere. in the end, people are always gonna be bias and give a bullshit score to teams they don't like or don't feel are deserving. showing how many standard deviations of bullshit they gave doesn't fix anything. its over, analyzing it over and over isn't gonna do anything to the result. all it's showing is the faults of the system that we already know.

the only statistic that matters is to go on stage and give it your 100%.
 

jvirk

New Member
Messages
889
Nikhil B. said:
trying to solve this problem using math isn't going to get us anywhere. in the end, people are always gonna be bias and give a bullshit score to teams they don't like or don't feel are deserving. showing how many standard deviations of bullshit they gave doesn't fix anything. its over, analyzing it over and over isn't gonna do anything to the result. all it's showing is the faults of the system that we already know.

the only statistic that matters is to go on stage and give it your 100%.
+1
 

Kapil

Wear Responsibly 2.0 || One for All...All for FAUJ
Messages
133
Badwal said:
SO, in conclusion, there will be no perfect judge, and no perfect comp., the end
100% agree. It's like trying to find the non existent holy grail.

Edit: I agree with Swi that we CAN make it better though, good points below.
 

yraparla

SwizzeeMusic.com
Messages
2,072
This site is a great resource for a lot of people including competition organizers. In fact a great deal of rubrics and general judging procedures have been gleaned from discussions here

This conversation is BEYOND Elite 8's system. Even if every judge judged every team there are still issues of consistent ranges and distributions and they affect every competition.

Math is a tool, it can help normalize some of the data but the bigger gain here is developing a framework, a common language, to explain to judges hey, one of you can't be giving scores from 50-100 while another is only in between 80-90. This type of common understanding is invaluable.
 

kinnell

*Account Deactivated*
Messages
2,159
Badwal said:
SO, in conclusion, there will be no perfect judge, and no perfect comp., the end
Truth. But that does not mean you should not strive to have perfect and fair judging nor should you not strive to host the perfect competition. It's when people strive for perfection do we see some great things manifest themselves. Would we have great competitions like BBC and Elite 8 if the competition organizers simply wanted to host a good competition and not strive to hold one of the best?

Nikhil B. said:
trying to solve this problem using math isn't going to get us anywhere. in the end, people are always gonna be bias and give a bullshit score to teams they don't like or don't feel are deserving. showing how many standard deviations of bullshit they gave doesn't fix anything. its over, analyzing it over and over isn't gonna do anything to the result. all it's showing is the faults of the system that we already know.

the only statistic that matters is to go on stage and give it your 100%.
Please see Swi's reply. I rarely give out more than +1s but Swi's reply deserves at least a +10. As Swi pointed out, the bigger picture here is not the conclusion that this system did not work or what faults it has. The bigger picture is understanding what happened and how things played out. And your "only statistics that matters" statement makes me giggle. Hehe :)

And Kapil, you are the man. (for your initial post)
 

Meistro

Asi Shounk Nu Karaiyaan Muschaan Kundiyaan...
Messages
2,065
Harmeet said:
US West got murked! haha ouch.
aww, I can hear the bitterness from your post. Its coo, maybe one day you will get invited.
 

nb0913

New Member
Messages
280
Swi said:
This site is a great resource for a lot of people including competition organizers. In fact a great deal of rubrics and general judging procedures have been gleaned from discussions here

This conversation is BEYOND Elite 8's system. Even if every judge judged every team there are still issues of consistent ranges and distributions and they affect every competition.

Math is a tool, it can help normalize some of the data but the bigger gain here is developing a framework, a common language, to explain to judges hey, one of you can't be giving scores from 50-100 while another is only in between 80-90. This type of common understanding is invaluable.
It seems as if you edited your original post where you destroyed what I said haha. I don't think I made the point I was trying to make clear at all..which I kinda see lol; I was pretty vauge. I agree with you that using statistics to show the deviations in judging is useful and obviously clarifies the numbers into a word form. I've taken statistics classes, I understand how they work, I understood the z-score and the standard deviation analysis that kinnell provided.

My point was that no matter how you represent the flaws of the system, using the statistics, it's not gonna change the fact that as people, we're gonna have bias. Telling someone that if you don't judge properly our standard deviations will be skewed, doesn't mean they will judge fairly. Yes it's good to try to normalize the understanding of what kind variations of judging should given. Obviously if one judge gives a 95, and another gives a 50, something clearly went wrong in understanding of the rubric. Statistics will show this pretty clearly when the mean is something ridiculously low. That was the point I was TRYING to make..guess I didn't do a good job first time around. Oops.

And kinnell, I was trying to be deep haha, guess it didn't work :(. Good job with AEG btw..definitely a sick performance.
 

yraparla

SwizzeeMusic.com
Messages
2,072
Nikhil B. said:
Swi said:
This site is a great resource for a lot of people including competition organizers. In fact a great deal of rubrics and general judging procedures have been gleaned from discussions here

This conversation is BEYOND Elite 8's system. Even if every judge judged every team there are still issues of consistent ranges and distributions and they affect every competition.

Math is a tool, it can help normalize some of the data but the bigger gain here is developing a framework, a common language, to explain to judges hey, one of you can't be giving scores from 50-100 while another is only in between 80-90. This type of common understanding is invaluable.
It seems as if you edited your original post where you destroyed what I said haha.

I understand how they work, I understood the z-score and the standard deviation analysis that kinnell provided.

My point was that no matter how you represent the flaws of the system, using the statistics, it's not gonna change the fact that as people, we're gonna have bias. Telling someone that if you don't judge properly our standard deviations will be skewed, doesn't mean they will judge fairly. Yes it's good to try to normalize the understanding of what kind variations of judging should given. Obviously if one judge gives a 95, and another gives a 50, something clearly went wrong in understanding of the rubric. Statistics will show this pretty clearly when the mean is something ridiculously low. That was the point I was TRYING to make..guess I didn't do a good job first time around. Oops.

And kinnell, I was trying to be deep haha, guess it didn't work :(. Good job with AEG btw..definitely a sick performance.
haha no I def understood and once I took a second to calm down and was like "wtf am I doing?" his points are straight

it's definitely fair to keep things in perspective that this won't be an end-all to judging issues. But I think what was bothering me was that kinnell saleem kapil and I (the people who are posting the most) obviously know the limitations and like random people just seemed to bandwagon on the "judging is flawed" comments which is just like unnecesary you know?

like we're just sharing our enjoyment of the topic of stats related to something else we love and there's no reason to like be negative about it, but I agree that perspective is always a positive thing (my post was more in response to people just +1 instead of being constructive like you) :)
 

Pam

New Member
Messages
59
Kapil said:
Some notes:

1) You can't say with 90% confidence you "know", you can only say that if you repeated this experiment with the same conditions 90% of the confidence intervals would contain the true mean. Someone else said that's not significant but it IS at an alpha of 10% (most common used is 5%, so in that case you're right).

2) Kinnel you are assuming that Bhangra judging as a whole is normally distributed, which it may not be. All judging may be skewed to the higher end (note no scores below 50 at all). Also, you can't use the Law of Large numbers or the Central limit theorem because your sample size isn't big enough for this example; you'd need many competitions' judging data. T-distr could be better since it's got fatter tails, but still as DF increases you approach the normal assumption.

3) If you were to mix game theory with statistical inference, but you would have to take into account probability of certain events specific to certain teams, thereby creating a more complex bayesian discrete distribution per judge on a team by team basis. Then you could compare distributions by assuming a hypothetical distribution, doing a siimulation of how he/she would judge at other comps, and compare the distributions vs other judges with a Chi Square Test due to large sample probability laws... I think....hah

Short version: Crapton of intricacies --> hard to actually model. You'd have to infer a judge by judge distribution, simulate, and compare if they are actually similar. One fact you missed out is that a judge may inadvertently make mistakes in #s due to subjectivity of writing down #s on the fly during judging, thereby giving us the "garbage-in, garbage-out" problem with the data; judges can't go back after they see other teams and "fix" the score.
lol, reminds me of my stat exam last semester...

Remy said:
Badwal said:
SO, in conclusion, there will be no perfect judge, and no perfect comp., the end
Well said
+1

i doubt there will ever be a competition where every single team/audience member/judge agrees on everything. when personal opinions and bias come into view there's no right or wrong answer.
 

zagreus

Active Member
Messages
1,473
I think at the end of the day, it's a bit of a flawed design to have teams judge each other. There's a lack of consistency, and i would prefer seeing a number of just qualified judges versus team judges. Team judges may be more biased due to their relationships with people on other teams, their respect for a team without critically looking at a performance, etc. Of course judges themselves may have a bias, but i would prefer a judge to watch all the performances versus a team judge that may see about 4 or 5 performances of the night.

But good show, and good job to the organizers for releasing the data like this.
 
Top