Login
You're viewing the front-end.social public feed.
  • Feb 1, 2026, 7:55 PM

    Qualitatively, I think there are things worth discussing. And it seems the authors agree, because there's 7 pages of it.

    They identified 4 "axes" of interaction with the AI that they think are explanatory. As they put it:

    - AI Interaction Time: The lack of significant speed-up in the AI condition can be explained by how some participants used AI. Several participants spent substantial time interacting with the AI assistant, spending up to 11 minutes composing AI queries in total
    - Query Types: The study participants varied between conceptual questions only, code generation only,
    and a mixture of conceptual, debugging, and code generation queries. Participants who focused on
    asking the AI assistant debugging questions or confirming their answer spent more time on the task
    - Encountering Errors: Participants in the control group (no AI) encountered more errors; these errors
    included both syntax errors and Trio errors (Figure 14). Encountering more errors and independently
    resolving errors likely improved the formation of Trio skills (note, Trio is the library that was the focus of the test problems)
    - Active Time: Using AI decreased the amount of active coding time. Time spent coding shifted to time spent interacting with AI and understanding AI generations (Figure 16).

    💬 1🔄 0⭐ 6

Replies

  • Feb 1, 2026, 7:59 PM

    I continue to find this frustrating, because they absolutely will not let go of notions like time to completion or rate of output as being meaningful, or even vital. But this is supposed to be a study about learning outcomes. I'm not an expert on the latest in learning/teaching research, but I don't think those are useful measures. Please, correct me if I'm wrong, though.

    They also have remarkably little to say about the control group in this section, and I think _that_ says a lot.

    💬 2🔄 2⭐ 20
  • Feb 1, 2026, 8:15 PM

    Still, they then identified 6 clusters of interaction patterns with the AI that correlated with performance and related in varying degrees to those axes.

    Personally, I think I would call it 5 patterns, plus a group that moved from 1 pattern to another. But maybe that's not an important distinction.

    They broadly fell into low-scoring and high-scoring groups. The low scoring group was:
    - AI delegation. Just having the chatbot generate all the code.
    - "Progressive AI Reliance". Or, starting with conceptual inquiry, and then later just having the chatbot generate everything.
    - Iterative AI Debugging. Just having the chatbot generate all the code, and then just showing the chatbot whatever errors resulted and instructing it to fix the problem.

    What I find *really* interesting here is that the group who started with what they call conceptual inquiry and then moved to delegation scored _30 percent_ lower on the quiz than the group who only engaged in conceptual inquiry.

    That is an ENORMOUS effect for a tiny intervention. They actually performed comparably the group that engaged in pure delegation the whole time. I don't see any discussion of this from the authors, and that also sucks.

    💬 1🔄 0⭐ 12
  • Feb 1, 2026, 8:16 PM

    I might have expected the initial approach that was more oriented around understanding would have some protective effect against the switch to a production orientation. I also wonder if it reflects a disengagement with the task? Apparently the authors don't share my curiosity.

    💬 1🔄 0⭐ 7
  • Feb 1, 2026, 8:29 PM

    Then there are the high-scoring patterns:

    - "Generation then comprehension". Generate code, and then ask followup questions about it.
    - "Hybrid code-explanation". Generate code and simultaneously ask for explanation.
    - "Conceptual inquiry". Don't generate code, just ask questions for understanding.

    The authors propose that "spending time" and "encountering errors" do a lot to explain the difference in quiz scores. The relative results from these groups make me doubt that. I suspect that the actual differentiator is having assumptions revealed and invalidated. The generate-then-follow-up pattern is the only one of the three that actually offers a chance to incorporate some result from a change into the explanation of the change. This group scored 16% - 19% higher than the other two, for nearly the same amounts of time spent on the task.

    💬 1🔄 0⭐ 3
  • Feb 1, 2026, 8:36 PM

    Finally, I think I'll leave you with some of the feedback given by the control group:

    - "This was a lot of fun but the recording aspect can be cumbersome on
    some systems and cause a little bit of anxiety especially when you can’t
    go back if you messed up the recording."

    - "I think I could have done much better if I could have accessed the coding tasks I did at part 2 during the quiz for reference, but I still tried my best. I ran out of time as the bug-finding questions
    were quite challenging for me."

    - "I spent too much time on this quiz, but that was due to my time management.
    Even if I hadn’t spent too much time on the first part, though, it still
    would have been a tight finish for me in the 30 minute window I think."

    To me, these read like stress. It's so disappointing that the study was designed in such a stressful way. Even moreso that the subject's stress doesn't seem to have been considered as a factor at all. That plus the tooling handicap of the control group make it impossible to draw the kind of conclusions that the authors and Anthropic seem to be doing.

    /end

    💬 1🔄 0⭐ 22
  • Feb 2, 2026, 1:08 AM

    Actually, one last thing.

    I don't think this study was well designed, but I don't want to go much farther than that.

    Some people have made something out of the paper not being peer reviewed. That's not a secret, though. This is arxiv, it's a prepublication host.

    Also, I glanced at the lead author's other work, and I get the impression that she's just not accustomed to working with human subjects. I think that's hubris, but not malice. It's just the standard attitude in tech that being good at computer touching qualifies one to do virtually anything else they want.

    💬 0🔄 6⭐ 26
  • 💬 0🔄 0⭐ 2