this post was submitted on 04 Jan 2024
177 points (90.4% liked)

Technology

59692 readers
3932 users here now

This is a most excellent place for technology news and articles.


Our Rules


  1. Follow the lemmy.world rules.
  2. Only tech related content.
  3. Be excellent to each another!
  4. Mod approved content bots can post up to 10 articles per day.
  5. Threads asking for personal tech support may be deleted.
  6. Politics threads may be removed.
  7. No memes allowed as posts, OK to post as comments.
  8. Only approved bots from the list below, to ask if your bot can be added please contact us.
  9. Check for duplicates before posting, duplicates may be removed

Approved Bots


founded 1 year ago
MODERATORS
 

ChatGPT bombs test on diagnosing kids’ medical cases with 83% error rate | It was bad at recognizing relationships and needs selective training, researchers say.::It was bad at recognizing relationships and needs selective training, researchers say.

you are viewing a single comment's thread
view the rest of the comments
[–] kromem@lemmy.world 16 points 10 months ago* (last edited 10 months ago)

This is a fucking terrible study.

They compare their results to a general diagnostic evaluation of GPT-4 which scored better and discuss it as relating to the fact it's a pediatric focus.

While largely glossing over the fact they are using GPT-3.5 instead.

GPT-3.5 sucks for any critical reasoning tasks, and this is a pretty worthless study not using the SotA or using best practices in prompting to actually reflect what a production grade deployment of a LLM for pediatric diagnostics would be.

And we really need to stop just spamming upvotes for stuff with little actual worth just because it's a negative headline about AI and that's all the jazz these days.