I get asked to advise on the board of a decent number of NSF proposals every year on human centered computer science questions. It is remarkable how similar they are. Lots of questions about measuring learning outcomes in CS or competency at programming and software work more broadly. I am not convinced there is a single unitary measure for this that we can quest after
@grimalkina I'm quite curious about the distinction between achievement measures vs aptitude measures! Thanks for that nugget, I have some reading to do.
@eqe we typically distinguish between the "three As": achievement (measuring what was accomplished in the past), ability (measuring a current capacity), aptitude (measuring potential... typically what we are trying to approximate but we use the previous two to do so. Wars fought, intellectually speaking, over how we do this, at)
Measuring skill and its development in a defensible and validated way is very, very hard. It is the type of work that requires you to bring in behavioral science, learning science, statistics, psychometrics, and multiple forms of testing theory. It is a set of topics I really enjoy but I really do see a lot of teams write proposals that sound unbelievably ambitious to me.
@grimalkina I agree with your assessment and approach. I saw too many instances of people—Management—assuming that a complex set of skills and understandings could be easily assessed or predicted.
@grimalkina I think it's interesting because every company has at least two processes that claim to measure skill and its development: hiring and performance evaluation. In my experience these are not scientific processes, but management will claim they are objective. They are not.
@welbog yeah, and much ink has been spilt on assessment in hiring! I do think we could make it better, some orgs have done so, but it's a real location for pseudoscience.
I think practitioners themselves are often drivers of pseudoscience in assessment though. Many "I'll know it when I see it" beliefs keep us using poor proxies and interview processes that we are familiar with and therefore think are accurate.
@welbog a tough thing I always think about here is the conditioning on a collider problem of this. We rarely directly observe the "would've been great but we rejected them" false negatives and so our perception that we're doing good selection is probably inflated.
@[email protected] @[email protected] I spoke to one manager who (with rare insight and humility in my experience) pointed out that if he couldn't be 100% confident in the accuracy of his interview process - which he was not - it was much more important for him to focus on "not hiring someone bad" than "not missing the best possible person". After all - hiring someone bad is an immediate negative impact, hiring someone who wasn't "the best" was at worst a small opportunity loss.
Having come to that conclusion, he then thought about what made a hire "bad" and immediately prioritised personality over technical skill, with the only immediate way to fail the interview process being to reveal yourself as an arsehole to the interviewers. You still might not get the job enough other candidates appeared to have better technical skills, but no amount of apparently technical skill would get you through if you (for example) answered questions a female interviewer asked you by speaking exclusively to their male colleage.
Apart from the fact I agreed with his priorities, it did give me pause for thought on a lot of the ideas around "objective hiring" (and I've worked in local government, so I've seen... things... done in an attempt to make hiring objective). Apart from the fact we're terrible at assessing competence, in many cases it feels like it shouldn't even be the most important factor.
@mavnn @welbog I love this story. Really important.
I think a lot about the kind of measurement injustice that is actually perpetuated when we act like a selection process is more predictive than it is. Like 1000 equally qualified learners will apply for 100 slots in a grad program let's say. In my view we should honestly do random selection once our evaluation criteria no longer reliabily discriminate.
@[email protected] @[email protected] Yes! The number of interview panels I've sat on where interviewers tried to find "the signal" to compare the remaining candidates when in reality our process just couldn't distinguish any further.
That's what I loved about this manager's approach. He liked to comment that if he told a prospective hire they'd be working with the "best of the best" he was revealing himself a lier and encouraging the arrogant to apply. While if he told them he would move heaven and earth to ensure they never had to work with an arsehole, it was something he could genuinely aim for with the tools available.
Of course, we did still want to assess competence/aptitude, and I don't really feel we did any better at that than anyone else. Nice team to work with though!
I think it is probably possible to develop better measures of multiple core competencies than we have, but those will be achievement measures, not aptitude measures (critical distinction!). I think the groundwork for this is sparse, besides concept inventories which are really suited to measuring undergrad threshold clearing, little validated work exists to measure problem solving directly in software work