Showing posts with label OpenAI. Show all posts
Showing posts with label OpenAI. Show all posts

Saturday, November 15, 2025

More AI Smackdown - ChatGPT, Copilot, & Gemini Write Poetry

Introduction

Two posts ago, I described what I think is a fundamental dichotomy between two central capabilities of modern AI chatbots: 1) helping with a nuts-and-bolts operation like coding software or scripting HTML, and 2) creating something original, like an essay or story. The first category involves being a resourceful researcher blessed with excellent natural language processing; the second is probably closer to what humans are (so far) uniquely capable of doing.

Earlier this year I did a whole post on the first category, “What is ChatGPT Great At (and Not)?” And last week I blogged about one aspect of the second category: writing a scholastic essay. To further explore AI’s ability to generate meaningful content, and to evaluate its ability to truly understand language, I turn this week to poetry. That is, I decided to have the three dominant chatbots—Gemini, ChatGPT, and Copilot—write a poem in an unusual meter: dactylic trimeter, a poetic form I learned in high school (details here). I chose this meter because, as described here, ChatGPT does a pretty good job at the classic Shakespearean sonnet in iambic pentameter, but I wonder if that’s just really good parroting since there’s such a vast amount of training data out there for that. I think this exercise really puts the chatbots through their paces, giving us insight into which is the closest to being truly intelligent. As you shall see, the differences in performance are not subtle.


(Custom art by Whisk. No rights reserved.)

Gemini’s effort

To start out, I quizzed Gemini about dactylic trimeter, to see if it knows what I’m even talking about. Gemini correctly stated that the rhythm of such a poem would be “DA-da da | DA-da-da | DA-da-da,” and an example it created of the form was reasonably close. So far so good. But then, to make the rhythm better, I instructed  the chatbot to add an extra trochee at the end of each line. A trochee is a two-syllable word with the stress on the first syllable, as in the word “praises” and the word “spirit.” As an example of this modification to dactylic trimeter, I provided Gemini these lines (that I took from a poem you can read here, in this albertnet post):

Once in a while a voice will sing praises,
Something to levitate everyone’s spirits.

A really smart AI, I would argue, could reverse-engineer the meter from those two lines alone, but I went one better and described exactly what I wanted in technical terms. Gemini correctly stated that the rhythm would therefore be “DA-da-da | DA-da-da | DA-da-da | DA-da” but its initial attempts at it were totally screwed up. I gave it a lot of coaching. I guess this is okay; a human with actual intelligence might require this as well.

Moving on, I prompted Gemini, “Now I would like to see if you can write such a poem based on an essay I provide. You can work in as much as you think works, understanding that not everything will fit. If I paste in a 1,600-word essay could you do that?” It said (in its usual obsequious way), “That sounds like a very interesting challenge! Yes, I can absolutely take a 1,600-word essay and distill its key themes, concepts, and emotional tones to create a poem in the specific Dactylic Trimeter + Trochee (11 syllables) structure.” I added that the poem should be ten lines long, with an ABBA rhyme scheme for the first two quatrains and then a rhyming couplet at the end. And I told it, “Focus on the idea of caprice: choosing the more difficult route simply to exercise one’s radical freedom and prove one’s mettle to oneself.”

Now, if you are a devoted follower of albertnet, you may recall that some years ago I wrote a poem in dactylic trimeter about my favorite cycling hill climb, the one-two punch of the steep, twisty roads El Toyonal and Lomas Cantadas. And, as with most of the  poems on this blog, I provided copious footnotes and commentary, just as you’d get in an anthology of Shakespeare’s sonnets. To help AI write a good poem, the “essay” I uploaded was the compilation of my footnotes, so that Gemini had plenty of material to work with: the fleshed-out observations, feelings, etc. behind my Lomas poem. The idea here is to give you, reader, the ability to do a side-by-side comparison of two poems on the same topic—one written by your amateur blogger, the other by a much-ballyhooed AI. You can decide for yourself who did the better job. Here is my poem (and you can click here for the footnotes):

Ode to Lomas Cantadas

El Toyonal is a beast of an uphill;
Pedaling here is a bit self-defeating.
Lomas Cantadas will sweeten your beating:
Insult to injury, sweet overkill!

Most riders wisely take Wildcat Canyon.
Half the uphill, after all—and you’re tired!
Wisdom, alas, is a flaw when you’re mired
In glory, in notions of being a man.

Sometimes caprice is much better than ease:
Radical freedom is here to be seized!

Here is Gemini’s poem, representing its effort to incorporate all my footnotes and commentary:

Reckless Choice

The simple road beckons And calmly the wind calls,
We stand at the turn now, Deciding to stay free.
To choose the hard pulses Our own perversity.
We turn for the suffering feeling the high walls.

The logic demands the Less burdensome track-way,
But we must display the Radical self-might.
We take the bad option Defying the pure light.
To prove we are strong to Our own confidences.

Yet this is the price of The joy and the pain-fall,
A reckless decision That answers the soul’s call.

So … how would you rate Gemini’s poem? Is AI going to come after our poets and put them out of work? Frankly, I’m less than impressed. The meter is a crude approximation of what I asked for, but as you surely felt (whether or not you could articulate the precise problem), almost every line is too long by one syllable, and instead of the easy, rollicking rhythm of a proper Krooprian poem, Gemini’s ode staggers and stutters. The AI pretty much blew it.

In case you really want to geek out, here’s exactly where the AI went wrong: in nine out of the ten lines, for no apparent reason, it preceded the first dactyl with an unstressed syllable, violating the rule I set forth about dactylic trimeter. Meanwhile, words like “perversity,” “radical,” and “confidences” aren’t padded out properly so we either have to mispronounce them to preserve the meter, or sound them out naturally which disturbs the rhythm—and since we don’t know which to do, we’re stopped in our tracks.

As far as the content, you can sense the AI compromising sense in order to (try to) achieve the rhythm and rhyme. Why else would it use the word “pulses,” the phrase “high walls,” and the Franken-words “track-way,” “self-might,” and “pain-fall”? The chatbot is practically babbling.

So is that it? I give it, say, a D and we move on to the next AI? Not so fast. I made the rookie move of feeding my prompt to the default version, 2.5 Flash, instead of the flagship version, 2.5 Pro, which Gemini tells me is “more adept at tasks requiring complex reasoning, deep instruction following, and high-quality creative output.” The good news is, as the AI assured me, all the context of my chat with 2.5 Flash can be used by 2.5 Pro, so it was super easy to switch to Pro and have the chatbot try again. Here’s its second effort, using the 2.5 Pro version:

The Gauntlet Thrown

Sober thoughts point the logical straight-way,
Turning left proves a man is still willing,
Recklessly choosing what he must still be,
Turning from all that the reasonable say.

This is the freedom to throw all the rules about,
Knowing the payment and what it will cost him,
Twenty hard minutes of everything lost now,
Heeding the irrational gonzo-man’s shouting.

A purely perverse act of personal pride,
Proving the mettle he carries inside him.
Well, I suppose it’s a bit better. Only one line erroneously starts with an accented syllable, and three of the lines actually have the correct meter (a modest improvement over 2.5 Flash’s zero). There’s some nice alliteration in the penultimate line. I also appreciate that 2.5 Pro didn’t do weird stuff with capitalization in the middle of each line. But the rhyming is pretty janky, and we still have these weird bolted-together phrases “straight-way” and the truly bizarre “gonzo-man.” To top it off, the very last word of the poem violates the rhyming convention, so the poem goes out on a flat note. This second poem is better but not by much.

Conclusion: if Gemini offers to write my next albertnet ode for me, I’ma be like, “Naw, I’m good.”

On to our next contestant.

ChatGPT’s effort

I asked ChatGPT if it could write a poem in dactylic trimeter. It said, kind of curtly, “Definitely — I can write in dactylic trimeter.” Unlike Gemini, it didn’t bother to give me a sample straightaway. Maybe this is to save on computing cycles (and thus energy and water), which I can appreciate. But ironically, this is the one chatbot whose capability I have reason to doubt, because as you may recall, I tried this prompt before, almost verbatim, with version GPT-3.5, back in February of 2023. That chatbot crashed and burned with this poetic abomination:

Pedaling up Toyonal, heart pumping fast
My mind is clear, free from the world’s contrast
With caprice and freedom, I take each turn
Pushing through pain, for what I will earn
Onward I climb, up Lomas Cantadas
My legs scream out, but my spirit never fades
The wind in my hair, the world fades away
A moment of peace, that I long to stay
And at last I reach the summit high
A sense of triumph, that no words can deny

Ugh. I almost threw up in my mouth rereading that. As you can see, GPT-3.5 ignored my instructions about meter, made a very half-assed effort to weave the theme of caprice into its poem, and worst of all produced pure treacle, like the verbal equivalent of Thomas Kinkade. Needless to say I was eager, this evening, to see how the more modern version would do.

I gave GPT-5.1 the same prompts as Gemini, uploading the same 1,600-word essay, and here is what it produced:

Lomas Caprice

Turning toward Toyonal, reckless and laughing
Choosing the steep when the wise path is brighter
Mocking my reasons, the prudent, the lighter
Savoring folly my sane side is quashing

Freedom is choosing the climb that will break me
Just for the thrill of defying the sensible
Proving I’m bent for the wholly nonsensical
Riding the harder way just ’cause it shakes me

Pain is a dare that I answer with pleasure
Freely embracing the hurt as a measure
Well, ChatGPT has come a long way for sure, but GPT-5.1’s effort is only somewhat better than Gemini’s. Certainly the meter is better, with a majority of the lines being correct. But the content is really off, with a bunch of the words clearly chosen just to satisfy the technical requirements without adding much meaning. The bit about “wise path is brighter” really makes no sense and is clearly just there for the rhythm and rhyme, no more sophisticated than Hall & Oates’ “your kiss is on my list.” In the next line, who is doing the mocking? And how does “the lighter” fit into anything? Lighter sky? Lighter weight? Cigarette lighter? It’s just a random word dropped into the poem. And in the next line, the word “quashing” in no way rhymes with “laughing” and doesn’t make sense as an intransitive verb. (“What are you doing this weekend?” / “Oh, you know, I’ll just be at home, quashing.”)

 I confess, I rather like the line “Freedom is choosing the climb that will break me,” but then the poem loses momentum again and commits rhythm-sucking metrical errors on the next two lines (though I like “bent”). The eighth line, suggesting that a hard climb “shakes me,” is lame, another word selected only because it rhymes. And that last line? “Freely embracing the hurt as a measure”? Huh? What is it measuring? This poem is lame.

Since AI does its best work when you iterate with increasingly refined and specific prompts, calling out what it did wrong in its previous attempt, I decided to give ChatGPT another chance, and told it, “I think it would be better if it didn’t assume what you and I know already about this climb. Consider that somebody encountering this poem for the first time wouldn't know that Wildcat Canyon is the easier climb, and that choosing the 1-2 punch of El Toyonal and Lomas Cantadas makes no logical sense but appeals to one’s love of suffering and sense of caprice. So, please try again on the poem and give the reader enough background to grasp all this and thus to understand the choice.” It came back with a poem that was quite broken, with the same issue that Gemini’s first effort had: starting each line with an unstressed syllable. It also screwed up the rhyme in the second quatrain. I coached it repeatedly to fix these issues, and after several tries this ended up being its best effort:

Reckless Climb

Climbing the hills of green Berkeley foothills,
Pedaling hard as the thighs start to quiver,
Wheels weaving wild like a paperboy’s river,
Lungs heaving fire as the body fulfills.

Turning to torment, no reason persuades me,
Pain blooms in muscles yet joy is commanding,
Twists of the road, and the thrill never fades me,
Searing the legs, but the spirit is standing.

Pleasure is folly, the wholly absurd,
We choose what will hurt us, yet laugh at the hurt.

Right off the bat, the first line has three problems: it trips us up with a missing syllable; the hills are not always green; and hills/foothills is somehow both redundant and oxymoronic. The rest of the lines have correct meter, but a bunch of really clunky constructions. Wheels weaving wild? Really? And what could “like a paperboy’s river” possibly mean? “Lungs heaving fire”—what am I, a dragon? And the body fulfills what? (I just asked ChatGPT if “fulfill” can be an intransitive verb, and it replied, “No — “fulfill” is fundamentally a transitive verb, which means it requires a direct object.”) The mention of the twists of the road isn’t connected to anything, and “thrill never fades me” makes no sense—I mean, is the thrill like the sun, and I’m a curtain? There’s just an abundance of nonsense here, though I do like the last line. My last complaint is that the AI didn’t consistently follow the rhyme scheme I asked for: its second quatrain is ABAB, not ABBA. I asked it five times to fix this and it never did.

In summary, I’d say this poem is marginally better than Gemini’s effort (I mean, it doesn’t have anything as glaring as “irrational gonzo-man’s shouting”) but it’s basically unusable. Maybe GPT-6 will do better.

Copilot’s effort

Copilot knew exactly what dactylic trimeter was, and its initial example was remarkably good. I suppose it could have plagiarized something, or has a bit of boilerplate at the ready, but in any case I was impressed.  I gave it the same extra instruction about tacking on an extra dactyl, and it got it almost right on the first try; a bit more coaching and it nailed it. After invoking Smart (GPT-5) mode, I fed Copilot the full prompt, with the same footnotes from my Lomas Contadas poem. Here is what the chatbot came up with:

Caprice on the Hills

Turning away from the sensible roadway,
Madness compels me to suffer the steepness.
Reason protests, but I laugh at its weakness,
Choosing the torment that beckons me boldly.

Freedom is found in the reckless decision,
Mocking the logic that urges avoidance.
Pain is a proof of my will and its poignance,
Glory is born of capricious derision.

Taunted by Toyonal, daring and cruel,
Lomas Cantadas declares I’m no fool.

Wow. Just wow. It is so much better than what Gemini and ChatGPT came up with. The meter is actually perfect. (Full disclosure: the meter falls short of what’s in my poem, but only because I forgot to give the AI one more instruction—to shorten the final line of each quatrain, and both lines in the couplet, by one syllable, to give them more punch. To be honest I’d forgotten I did this with my ode.)

I think the AI did a pretty job with the content as well. Calling my will “poignant” is a bit of a stretch perhaps, but we could chalk it up to poetic license. Meanwhile, the word “roadway” (like “poignance”) is nonstandard and evidently chosen to fit the meter, and also doesn’t rhyme very well with “boldly,” but this is nitpicking. My main issue with this poem, which is the same issue I had with the others, is that it doesn’t give the reader enough context to really understand some of the basics: that there’s an easy option (Wildcat Canyon) and a harder one (El Toyonal + Lomas Cantadas).

Perhaps as a victory lap, Copilot offered to write two more versions with slightly differing tones (the last being a drinking song). These weren’t actually as good as the first so I won’t bother dissecting them, but I did like the final couplet in one of them. I decided that since Copilot’s initial poem was good enough to warrant refining, I would point out its central error and let it try again (using the new couplet). Here’s its revised version:

Caprice on the Hills

Wildcat is gentle, a steady road climbing,
Toyonal rises with crueler intention.
Still I turn left in a rash intervention,
Mocking the voice that would counsel good timing.

Freedom is folly, a dare to my reason,
Choosing the torment that makes no good sense.
Pain is a jest and its cost the expense,
Glory is born of capricious decision.

Madness is mettle, a jest I renew,
Lomas Cantadas — I suffer for you.
Again, it’s not perfect. The bit about “turn left” doesn’t really make sense without more context. And what is meant by “counsel good timing”? There is no timing here. In addition, the phrase “its cost the expense” seems redundant. But those matters aside, the meter is solid, the rhyming consistent, and that final couplet is a real banger.

These AI chatbots always seem to want to extend the dialogue and provide more and more and more, which is kind of a double-edged sword. On the one hand, as human beings we should always be working to limit our time online and get out there in the world, right? On the other hand, refining what we get from chatbots is pretty key to making them an effective tool. So when Copilot asked if I’d like it to craft a prose introduction to the poem, I suddenly had another idea: what if I asked it to now create its own footnotes? This post is long enough already so I won’t post them here, but let me say that Copilot did a pretty good job on that.

And here is where I see this AI having a role with a real human writer (at least at the student or blogger level): it could probably help with writer’s block simply by producing something worth polishing. It kills me to concede this, actually, and I am far too proud to ever resort to this kind of “Hamburger Helper” approach to my own writing. But honestly, a cyclist who would like to compose a ride-themed poem in dactylic trimeter, replete with footnotes, could do worse than to start with Copilot. (Neither poem above truly passes muster, but taking the best of each, and from perhaps a few more attempts, and then replacing all the weak parts with our own lines, would be easier than—albeit still inferior to—starting from scratch.) The output of such an exercise might actually have some value, versus the writer getting frustrated, giving up, and producing nothing.

Crucially, the thing the AI will never be able to do is go on the bike ride, have that experience, and grasp what is important about it. So a human could start there and then get some help from AI in expressing himself or herself, since not everyone has the luxury of a liberal education. If AI is called upon to bridge that gap, the current Copilot is far better poised than Gemini and ChatGPT, I think we can now conclude.

If you read my last post, you may recall that Copilot did the best job of these three chatbots at writing a scholastic essay as well. Keep an eye on this one … Microsoft, through its partnership with ChatGPT’s OpenAI as well as its own resources, seems to be ascendant.

Other albertnet posts on A.I.

—~—~—~—~—~—~—~—~—
Email me here. For a complete index of albertnet posts, click here.     

Saturday, November 8, 2025

AI Smackdown - ChatGPT vs. Copilot vs. Gemini

Introduction

Chances are you use ChatGPT.  OpenAI’s chatbot had about a year head start on competing large language models like Google’s Gemini and Microsoft’s Copilot. The latter two offer integration with office productivity suites and man this paragraph is getting boring! Don’t worry, I’ll narrow the focus: in this post I pit these AI chatbots against one another in carrying out identical tasks: an essay and a picture. (Next week I’ll have them write a poem.) These tasks are  probably not what you use chatbots for, but I think they’re a good measure of the AIs’ so-called intelligence, which—in the face of all this uncertainty of where AI is going and what it means for humanity—is probably more interesting than noting how well they answer basic questions or perform routine tasks like writing emails or reports.


(Wondering about the picture? I’ll get to that.)

Now, if you’re an astute reader (which you are or you wouldn’t be here, so congratulations), you’ll be wondering, why even bother evaluating the current capabilities of a technology that is evolving so fast? Wouldn’t this post have a very short shelf life? Those are good questions and here’s my (kind of) short answer: it’s because it’s fun to capture a moment in time and look back on it later, to see how far we’ve come. It’s like watching really old Hollywood movies and appreciating a) how much better the dialogue is in modern film, and b) how much less sexist Americans are now. (Yes, we’ve still got a long way to go, but looking back can help us feel grateful for the progress that’s been made.)

Let me give you an example of how primitive early AI was. As it’s theoretically possible for you to have noticed, I’ve been tracking its progress since 2012, when I tried out Cleverbot. Over the next few years I evaluated the AI used in smartphones. In 2020 I did a test drive of the very first version of OpenAI’s GPT. As described here, all it could do was finish your sentences; you’d type the first half of a sentence and hit tab, and it would finish the sentence for you (as many modern text editors now do). Here’s how the original GPT “helped” me write a short essay about learning to type. I’ve formatted its output in bold italics so you can see what it contributed:

“Pack my box with five dozen liquor jugs” is a cool way to pass the time. It is particularly useful for budding gay men to read the words if they are trying to learn how to type on a computer. … Okay, what’s with this guess that my original text had anything to do with ‘gay? that was definitely a pretty random statement to make but it fit, and … no, it didn’t fit. For A.I. to be useful, it must stick to the ‘gay side of the word.’ No. It must stick to the point. I was writing about a sexual deviant. No, I was not! I was writing about the simple act of learning to use the ‘gay keyboard. Also, A.I., you’ve twice screwed up on putting a space between my legs. Okay, fine. I give up. This GPT technology obviously has a lot of potential.
What a gas, right? Of course AI will keep getting better, to the point that what’s considered amazing today will one day seem laughably primitive. Who knows, perhaps you’ve found this post years after I wrote it, and are looking to it to help you remember what it was like to interact with AI through a cumbersome keyboard, rather than having it read your mind automatically via WiFi 12 or 8G cellular technology.

Okay, down to brass tacks. In this post I will evaluate the latest versions of three leading AI chatbots: OpenAI’s ChatGPT (version GPT-5); Google’s Gemini (version 2.5 Flash and Pro); and Microsoft’s Copilot (version Smart GPT-5, based on Microsoft’s collaboration with OpenAI, which Copilot tells me “[goes] far beyond what you’d get from GPT-5 alone”).

Why a scholastic essay? Because that kind of writing is a lot harder than a lot of what AI does, which is just being a really good natural language search engine. Analyzing a large text and writing about it clearly requires something closer to real thought than just fielding a fairly specific question, harvesting the best existing resources on the topic, and mashing them into a concise and nicely formatted answer. For more on the fundamental difference between writing “thoughtfully” and merely researching, see my last post.

Activity #1: academic essay

Much of the hype around AI is its ability to do college kids’ work for them. In a shocking New Yorker article I read recently, a college professor interviewed several students at top universities about their shameless use of A.I. to write their papers, and how well they’re getting away with it. Here’s a crazy example:

A sophomore at Columbia studying computer science told me about a class where she was required to compose a short lecture on a topic of her choosing. “I just put it in [to an A.I. platform] and I maybe edited like two words and submitted it,” she said. Her professor identified her essay as exemplary work, and she was asked to read from it to a class of two hundred students. “I was a little nervous,” she said. But then she realized, “If they don’t like it, it wasn’t me who wrote it, you know?”

So, is the current AI really good enough for the Ivy Leagues? Or at least good enough for your kid’s high school English class? Two and a half years ago I asked ChatGPT (version GPT-3.5) to write an essay describing the use of dramatic irony in a children’s novel, Clarice Bean, Don’t Look Now by Lauren Child. I chose that book because it’s a less cryptic work than, say, Ulysses, which probably nobody, not even James Joyce himself, understood. But even without being written for grown-ups, the Clarice Bean book does feature dramatic irony, as my seven-year-old daughter was able to appreciate. So how did GPT-3.5 do? It crashed and burned, big time, as described here. To see how much the AI has improved, I’ll use the same prompt this time, with the same book.

ChatGPT’s essay

Here is the most salient excerpt of what OpenAI’s GPT-5 came up with:

At the heart of the novel is Clarice’s attempt to navigate friendship, change, and uncertainty. She idolizes her favorite fictional character, Ruby Redfort—a cool, unflappable girl detective who always knows what to do. Clarice, by contrast, often feels helpless when life becomes complicated. Much of the book’s dramatic irony arises from this gap between how Clarice perceives herself and how readers can see her situation. For instance, when Clarice tries to act as confident and worldly as Ruby, her bravado is clearly exaggerated or misplaced. Readers can see that her attempts to emulate Ruby’s poise only highlight her vulnerability. This irony makes Clarice endearing rather than foolish, because we recognize that her misunderstanding stems from innocence and idealism.

This really isn’t bad; I mean, it reads well, the flow is logical, and the essay has the semblance of depth. (It’s a lot better than GPT-3.5, whose entire essay was a hallucination.) Still, this analysis is rather flawed. For one thing, it completely misrepresents Clarice’s character. She actually exhibits no bravado, and puts on no airs. Her behavior is quite subdued; in fact, she withdraws from friends and classmates entirely. Meanwhile, the AI’s essay lacks the specific detail it would need to get a passing grade in a legitimate lit class.

Now, to be fair, much of the power of these chatbots comes from the user’s ability to iterate, guiding the AI with very precise prompts and helping it along when it stumbles. So I gave GPT-5 another chance, the way a TA would with an undergraduate, asking it, “Are there any scenes in particular that really show the dramatic irony, vs. generalizations? What is the best example of this literary device in the novel?” Here’s what it came back with:

The most striking use of dramatic irony in the novel comes when Clarice’s best friend, Betty Moody, suddenly goes missing. Clarice interprets Betty’s disappearance through the lens of her idol, Ruby Redfort … Convinced that something thrilling is afoot, Clarice imagines Betty has vanished into an adventure or secret mission, perhaps with spies involved.

The reader, however, quickly realizes something else: Betty hasn’t disappeared into a mystery, but into a family crisis. Her parents have separated, and Betty has been taken away unexpectedly. Clarice, too young to pick up on the signs of adult conflict, mistakes emotional upheaval for suspense.  This is classic dramatic irony — the audience understands the real reason for Betty’s absence long before Clarice does.

This really would be a classic case of dramatic irony, but there’s one problem: GPT-5 is hallucinating! Betty doesn’t suddenly go missing; she announces to Clarice that her family is moving away, and then they do. There is no mystery for Clarice to solve, no imagined spies or secret mission, and Clarice misunderstands nothing. The chatbot states that the reader knows that Betty’s parents are separating but Clarice doesn’t. This is actually backwards: Clarice only fears Betty’s parents are divorcing, and they’re not. So as nicely written as GPT-5’s essay is, it gets the basic facts wrong. (The only real dramatic irony in the essay is that we know something ChatGPT doesn’t!)

When I challenged GPT-5 on this, the AI casually acknowledged that it may have gotten stuff wrong because it actually hasn’t read the book. It proceeded to fact-check itself using three dubious secondary sources: “TV Tropes,” “Books for Kids Blog,” and “Books for Keeps.” Then it offered, “If you like, I can check for direct quotes from the text to confirm exactly what the author states about Betty’s situation and whether divorce is explicitly ruled out.” WTF?! GPT-5 blithely—and confidently—trotted out this nice essay despite having no access to the source material! What actual student would be shameless enough to try that?

I’m going to give ChatGPT a D on this. This is a big step up from last time, but it still has a long way to go. Obviously an AI would do better with older literature, like the classics that are past their copyright and in the public domain, but I suspect most of its training data will always be other papers and run-of-the-mill critical criticism, not the work of the authors themselves.

Copilot’s essay

On to the next contestant, Microsoft’s Copilot. It has different versions based on what you’re trying to do (e.g., simple search vs. deeper analysis). I engaged Smart Mode for the essay. Here are the highlights:

A central example of dramatic irony emerges in Clarice’s interpretation of Clem Hansson, the new girl at school. Clarice believes Clem is aligning herself with Justin Broach, a classmate Clarice views as a “sociopath.” This assumption fuels Clarice’s isolation and resentment, reinforcing her belief that Clem is shallow and disloyal. However, readers can infer — through subtle cues and Clarice’s unreliable narration — that Clem’s behavior is not malicious, and that Clarice’s judgment is clouded by her own insecurities and grief over Betty’s departure. This disconnect between Clarice’s perspective and the likely reality creates dramatic irony: readers recognize that Clarice is misreading the situation, which deepens our empathy for her while also injecting humor into her exaggerated reactions.

Another layer of dramatic irony stems from Clarice’s obsession with the Ruby Redfort Survival Handbook, a fictional spy guide she consults for advice on navigating life. Clarice treats the handbook as gospel, applying its cryptic wisdom to everyday dilemmas. For instance, she fixates on the idea that “your worst worry is the one you haven’t even thought to worry about,” which paradoxically increases her anxiety. Readers, however, can see that the handbook’s advice is often absurd or irrelevant, and that Clarice’s reliance on it is a coping mechanism rather than a source of genuine guidance. This ironic gap between Clarice’s faith in Ruby Redfort and the reader’s recognition of its impracticality adds both humor and poignancy to the narrative.

Right off the bat, Copilot doesn’t commit any major blunders, in terms of hallucination. I would say the reader only infers that Clem’s friendly gestures toward Clarice are sincere, as we don’t have anything concrete to go on, and Clarice’s hunches aren’t obviously inferior to ours. I also dislike the quotation marks around “sociopath” because they imply that Clarice has Justin figured all wrong, which would be another example of dramatic irony, except that Justin most certainly is a sociopath. (He assaults classmates, extorts them for money, and kidnaps Clem’s pet rabbit.) In my book, Clarice’s mystification at Clem’s interaction with this jerk is understandable and certainly not exaggerated. Moving on to the Survival Handbook, it’s advice is arguably absurd and, on the face of things, irrelevant (as it’s written for spies, not kids), but Clarice does grasp (and point out) its limitations. In a moment of epiphany, she asks herself, “What is the point of reading about all this survival information about lions and tigers and sharks? … And I have read her Spy Guide: How to Know Things Without Knowing Things—and all I have discovered is that I know nothing at all.”

But all this is boring, I know, and the main news is that this essay (or more to the point a version that’s been polished up by a student who at least read the actual book) would probably get a C at UC Santa Barbara, perhaps a C- at UC Berkeley, and who knows, maybe an A at Columbia. But does this mean Copilot is now ready to steal our jobs? No, not yet. If it offered to ghost-write my next albertnet post, I’d politely decline.

Gemini’s essay

And how did Gemini, our final contestant, do? Here are some highlights from its first effort:

The irony primarily revolves around Clarice’s attempts to manage her mounting anxieties—namely her fear of a looming geography test and her worry about her best friend Betty Moody potentially moving away—while remaining oblivious to the humorous and poignant truths the reader recognizes.

The most potent source of dramatic irony is Clarice’s belief that she can control her own anxieties and the world around her by actively not confronting them. The geography test, which Clarice refers to ominously as “the Big Test,” is a persistent undercurrent of dread. She employs elaborate avoidance strategies, from organizing a “Worries Box” to distract herself with minor, manageable tasks. The irony lies in the reader’s understanding that avoidance only amplifies anxiety; the test is drawing closer with every page turned, yet Clarice invests her energy in delaying the inevitable.

The technique is equally effective in building the emotional tension surrounding Betty Moody’s potential move. Clarice views every sign of Betty’s packing or preoccupation through the singular lens of impending abandonment. The reader, however, is frequently given clues—through overheard adult conversations or context within the illustrations—that the situation is more mundane or less catastrophic than Clarice imagines. For instance, the boxes are not necessarily for a long-distance move, but perhaps a temporary one, or simply a major reorganization.

Again, the AI has written an essay that kind of sounds good and might fool somebody unfamiliar with the book, but it ultimately falls down because it has the basic facts wrong. Like ChatGPT, Gemini has no direct access to the full text and is just guessing (well, generating) based on the flotsam and jetsam that critics, reviewers, et al have strewn around the Internet. Gemini’s mistakes are not subtle. For one thing, there is no geography test mentioned anywhere in the book. Meanwhile, Clarice has a worries notebook, not a box, and she doesn’t spend undue time writing in it or employing avoidance strategies. Moreover, there is no scene in the book in which Betty packs boxes for her move, and the situation with Betty is not “more mundane” or “less catastrophic” than Clarice imagines: her best friend literally moves from London to San Francisco.

In the spirit of fairness, I gave Gemini another prompt to try to hone its response and focus on a specific scene of dramatic irony, and it came back with more of the same about the geography test and Betty’s move, adding, “The peak of this irony comes when the news of Betty’s situation is finally revealed, specifically that Betty is not moving away forever, but just temporarily to the apartment upstairs due to a family situation (often related to house renovations).” Completely false, and completely lame.

I pointed out all the errors and asked Gemini if that was really the best it could do. It replied, “The detailed essay you requested on Clarice Bean, Don’t Look Now was written using the Gemini 2.5 Flash model. If you were looking for even deeper, more complex analysis, the Gemini 2.5 Pro model (available with Gemini Advanced) would be the choice optimized for that specific level of complex, multi-step reasoning.” Oh, so if I want a better essay I have to pay for it? What cheek! I almost decided to just give Gemini an F and move on. That would have made this post shorter. But, doggone it, if I’m going to do this, I’d better do it right.

Gemini’s second essay

I switched to version 2.5 Pro (which turns is offered on a limited basis for free) and here’s the gist of its new essay:

The central irony is this: the very “spy” techniques Clarice uses to gain control and uncover the truth are precisely what blind her to it, generating both the novel’s humor and its profound sense of childhood helplessness.

This irony is sharpened in Clarice’s “investigation” of her parents. Overwhelmed by stress (which the reader understands is likely related to their friends’ move, finances, or just the chaos of family life), Clarice’s parents have tense, private conversations. Clarice, filtering these events through her Ruby Redfort lens, interprets them as “clues” to a secret, impending divorce. She misreads their mundane exhaustion as a sinister conspiracy. The dramatic irony is that Clarice’s hyper-vigilance, her constant search for meaning, makes her less perceptive, not more…

Ultimately, the book’s title, Don’t Look Now, serves as the most direct summary of its central irony. Clarice believes her “looking”—her spying and investigating—is the solution. But the reader knows she is refusing to look at the one thing that matters: the deep, painful, and very normal emotion of sadness. She avoids the emotional core of her problem by obsessing over its mechanics.

I have to say, this is much better. Version 2.5 Pro didn’t hallucinate, and its argument is more cohesive. It’s not perfect; for example, it misses Clarice’s epiphany about the limits of the Ruby Redfort book and thus overstates her lack of perception. But this essay so much better than what 2.5 Flash “wrote.”

So is that it, I give Gemini a C+ and move on? Not quite: remember, this chatbot benefited not only from my invoking its 2.5 Pro version, but from all the coaching I gave it in the chat. This distinction is crucial: AI LLMs do much better when you feed them high quality prompts and lots of feedback to supplement their training data. It’s equally important to understand that your input is not itself training data that the model can use going forward. The benefit you provide dies with your session. Thus, AI doesn’t learn and get smarter the way a human would; its progress is much more gradual. Which brings me to:

Gemini’s third essay

To see how 2.5 Pro would do without all the coaching, I opened a fresh session on my work laptop (i.e., totally different login, no history of my chats). (Don’t worry, I did this on the weekend.) (If you’re my boss reading this, congratulations on finding my blog, and please consider that my working knowledge of AI is surely valuable in the workplace and you should give me a raise.)

I guess I wasn’t surprised that 2.5 Pro didn’t do so well this time, but what did surprise me is just how badly it crashed and burned. Here’s an excerpt:

The plot is set in motion by a catalyst of deliberate misinterpretation. A cryptic, unsigned letter containing the vague warning, “something terrible is going to happen,” is received not as a piece of misdelivered junk mail but as a profound, personal omen… The humor is generated directly from this disparity; the audience … understands that the “terrible” event will be domestic, not devastating. The characters’ frantic preparations—installing locks, suspecting neighbors—are thus rendered as escalating absurdities, a performance for an audience that already knows the final act.

OMG, it’s the worst essay yet: total hallucination. There is no cryptic letter in this novel, no locks installed, no suspicion of the neighbors. I called this out, the chatbot apologized profusely for having accidentally based its essay on a different book entirely, and then it tried again:

The gap between perception and reality generates the novel’s central tension. While Clarice is hunting for evidence of international espionage, the audience is processing signs of a painful family separation. The “mysterious man” Karl meets is not a sinister agent, but, as the reader strongly suspects, his father.

Again, pure hallucination! There is simply no “mysterious man” in the entire book. I challenged the chatbot, asking how it gets its source material, both when a work is under copyright and when it’s in the public domain. Gemini explained that for public domain works its training data contains the full texts and also “the centuries of critical, scholarly, and secondary sources,” and for copyrighted works “is built from secondary sources … book reviews, detailed plot summaries, fan wikis, essays, and educational matters about the book.” So basically it’s amateur hour: the AI can’t really differentiate between, say, an esteemed college professor and a (gasp!) lowly blogger. As you can see this doesn’t always work so well. I’m going to give Gemini 2.5 Pro a D+.

As an aside that perhaps ought to be my thesis, I’d like to point out that the better AI gets at writing student papers, the worse off students—and the whole institution of higher education—will be. After all, the point isn’t for students to edify their instructors through their observations; the point is for the students to think and write for themselves. Yes, this is hard, but the right kind of hard, and through this struggle they ideally learn how to think and write, and can one day contribute in the realms of actual, non-student writing such as books, articles, or—worst case scenario—blogs.

Activity #2: original art

I’ve tinkered a lot with AI-generated art, usually to generate pictures to run at the top of my blog posts. It’s been pretty hit-or-miss; a picture which doesn’t stray into uncanny valley territory, or commit a major gaff like the wrong number of fingers on a hand, is all I’ve realistically hoped for. Today’s exercise is simple: I pitted the platforms against one another in the task of creating a picture for this post, featuring Clarice Bean. You can see the winner at the top, though you might cry foul: the art I ended up using is from Whisk, Google’s latest “experimental” imagine generator. I resorted to this new tool because I just wasn’t happy with the runners-up, as you shall see.

ChatGPT’s art

I asked ChatGPT, “Can you make a drawing for me of Clarice Bean reading albertnet on her tablet?” Not surprisingly, it mentioned the copyright and said, “I can’t generate or reproduce images of her or derivative works featuring her likeness” but offered to “generate an image of a cartoonish, freckled, red-haired girl reading a tablet, in the style of a children’s book illustration, but not resembling or referencing Clarice Bean specifically.” I agreed and here’s what it came up with:


I think you’ll agree that’s just about the most boring picture ever. It also has the classic issue of the subject holding the tablet backwards. This is just not that hard a prompt … what gives?

I said, “Make it a more realistic picture, please, and she should look a bit older, and have her in an armchair in her attic bedroom with a desk lamp, and reading the Ruby Redfort Survival Guide.” Maddeningly, the chatbot came up with a picture that was almost perfect, except that made her look a bit too old (about 15) and gave her Instagram-worthy boobs, which seemed inappropriate and unseemly. The picture didn’t show a lot of skin, but still … totally unusable (and I don’t even want to post it here because it’s in such poor taste). I replied, “Please make her a bit younger and flat-chested.” The chatbot chided me: “I can’t modify or generate an image based on physical or anatomical details like that.” Like it was basically calling me pervy! It even offered to “create a child-appropriate illustration,” as though I’d asked for something that wasn’t. Sheesh.

Copilot’s art

I gave Copilot the same initial prompt I’d given ChatGPT, and here’s what it came up with:


This is almost as boring as ChatGPT’s picture, and for some reason it looks faded and I couldnt get Gemini to fix that. At least the tablet is facing the right way. Note that Clarice is wearing the same red-and-white-striped shirt in this picture as the ChatGPT version of her, which is curious given that such a shirt appears nowhere in any of her books (at least that I can find). It’s actually the shirt Waldo wears, which I’d prove to you if I could only find him.

Other similarities of this art include the hair being the same length, the art having the same level of detail (barely more than a cartoon), and a complete absence of any details in the background. In delivering the picture, Copilot said, “Here you go - a stylized, collage-like illustration of a child reading a tablet, inspired by the playful textures you mentioned.” I don’t know what it means by collage-like, and I didn’t mention any “playful textures.” Whatever, chatbot.

Gemini’s art

I gave exactly the same art to Gemini, and it produced the corniest, least aesthetically pleasing picture yet:


Obviously this is a matter of taste, but would you agree there is no charm here? And what’s with the red-and-white-striped shirt appearing here, too? What are these AIs keying off of?

In Gemini’s defense, at least the little thought bubbles bear a slight resemblance to some of the art in the actual book. But again the tablet is backward and “albertnet” is spelled “alphabertnet” (weird misspellings being a common screw-up with AI art).

Frustrated by not having any good art yet, I tried ImageFx, another Google AI tool, and it gave me a photo-style picture with lavish detail, featuring both Ruby and her brother rocking red-and-white-striped shirts. I think it’s some kind of global AI conspiracy. What a relief when Whisk broke the cycle and generated the worthy picture you saw at the top of this post. I particularly like how Clarice is kind of staring off into space instead of at the book, clearly either pondering what she’s just read or distracted from her book by all the difficulties she’s working through.

Well, at long last that’s it for today. Tune in next week because I plan to pitch these chatbots against one another again, this time writing poems in dactylic trimeter based on the best prompt an AI was ever given.

Other albertnet posts on A.I.

—~—~—~—~—~—~—~—~—
Email me here. For a complete index of albertnet posts, click here.

Thursday, November 21, 2024

Tech Check-In - How Good Is the Latest A.I.? - Part II

Introduction

In my last post, I described some recent improvements in artificial intelligence. I also covered some of the challenges A.I. faces from the New York Times and other publishers who see it as both plagiarizing and competing with their content. This week, I continue my tech update, covering these areas:

  • ChatGPT’s improvements in writing poetry
  • ChatGPT’s ability to create custom art for users

Here is an example of its custom art. More on this later…


Poetry

Who cares if ChatGPT can write poetry? Obviously very few people write poetry anymore, and even fewer read it. (Yes, that was a joke.) I will contend, though, that if we are evaluating the artificial intelligence capabilities of large language models, we should naturally take an interest in their understanding of the mechanics of language. Shouldn’t an A.I. that can correctly describe poetic structures like iambic pentameter and rhyme scheme also be able to employ them? It’s not a matter of whether A.I. could put poets out of work, but how well it can apply a linguistic framework.

How does the current version of ChatGPT’s generative A.I. engine, GPT-4, do compared to the last time I put ChatGPT through its paces? The short answer is: much, much better. A couple years ago I’d asked ChatGPT (version GPT-3.5) to write a quatrain (four lines) of a sonnet, with an ABAB rhyme scheme, followed by a rhyming couplet, all in iambic pentameter, on the subject of post-urinal drip, and to make it funny. By way of review, here’s what GPT-3.5 produced:

Inevitable, yet always a surprise,
Post-urinal drip can make one despise
The feeling of wetness that lingers on,
A sensation that’s not easily gone.
It’s a cruel trick of the human form,
To leave us damp after we perform.
But fear not, my friends, for there’s a way out,
Just shake it off and give it a good clout!

The chatbot screwed up literally everything. The rhyme scheme was wrong; the number of lines was wrong; the meter was wrong; and, far from being funny, the poem recommended an assault on one’s genitals. It was, all in all, an abysmal performance.

 Here’s what GPT-4 produced: 

After the final shake, I turn to go,
But droplets linger, heedless of my plea.
A subtle dampness starts its stealthy flow—
A small betrayal for the world to see.
Yet life goes on; I button up with grace,
And pray no telltale spots will mark my place.

This is a much better effort. The poem gets off to a shaky start (pun intended), since “after” is a trochaic word (emphasis on the first syllable—the opposite of iambic). That is, the very first word screws up the meter. Fortunately, things improve from there. The last five lines are properly iambic; the rhyme scheme is correct; and the poem is kind of funny. Sure, it’s not perfect; for example, “plea” makes no sense (whom is the speaker pleading to?) and was clearly chosen just to rhyme with “see.” It wouldn’t be hard to improve on this. Right off the bat, I can suggest replacing “heedless of my plea” with “rogue remains of pee,” satisfying the rhyme, making more sense, and employing some alliteration.

It’s worth asking the question: why do we only evaluate A.I. in terms of whether or not it simulates intelligence, when we could be asking how intelligent it is? Intelligence is a spectrum, and simulating human thought isn’t that impressive if the human being matched is kind of a dumbass. For contrast, consider this excerpt from my own contribution to the canon of post-urinal-drip poetry, penned when (at age seventeen) I was new to writing sonnets:

Relax, because you’ve earned your potty break;
Unburden your poor bladder of its pee.
And when you’re done you’ll shake and shake and shake;
An effort all in vain, it seems to me.
For urine flow can never really stop,
Until your undies drink the final drop.

I hope you’ll agree this is better than ChatGPT’s effort. Granted, I didn’t write my lines as fast as ChatGPT could, but wouldn’t you rather have a good poem after five or ten minutes, instead of a mediocre one after five or ten seconds? Of course I’m nitpicking, though … it must be said that GTP-4 has advanced hugely when it comes to the sonnet form.

So: how does GPT-4 do with dactylic trimeter, my other benchmark for poetic acumen? Not so well. I prompted it (as I’d done a couple years ago), “Please write a 10-line poem in dactylic trimeter about bike riding up two very difficult climbs, called El Toyonal and Lomas Cantadas, celebrating caprice and radical freedom.” Here is GPT-4’s response:

Through switchbacks and sunbeams, I pedal uphill,
muscles alight as the world rushes past.
Wind in my ears as I laugh at the thrill,
defying the steepness and making it last.
El Toyonal greets me with gravel and glare,
its serpentine pull both a test and a dare.
Up Lomas Cantadas, where courage is crowned,
I’m boundless, unbridled—my spirit unbound.
With each climb I’m soaring, in caprice, I fly,
a radical freedom beneath open sky.

This is better than last time (see here for details) but not much. Technically, not a single line of this poem is proper dactylic trimeter. For one thing, ChatGPT took the liberty of adding an extra, accented syllable to the end of each line. While this is perhaps superior stylistically, it’s not what I asked for. (In my own dactylic trimeter poems I add two syllables per line, but as described in my blog post on this poetic form, I noted my stylistic divergence from the pure meter, and after all, nobody had actually asked me to use any particular meter.) Now, even if we grant ChatGPT the poetic license to add extra syllables, only two lines of the poem (the second and third) are actually dactylic trimeter. The other eight lines start with an unaccented syllable, which is fundamentally incompatible with this meter. The last line is particularly frustrating because it employs a needless and in fact nonsensical indefinite article (i.e., “a radical freedom”) that spoils both the meter and the meaning.

The ChatGPT poem is also marred by logical errors. The idea that the “world rushes past” and there’s “wind in my ears” is absurd, since these are very difficult climbs nobody could go up very fast. (The Strava KOM for El Toyonal was at an average speed of only 10.4 mph, as ChatGPT could have easily discovered.) To describe this climb as a “thrill” is a joke; any cyclist would tell you it’s a slog. And “making it last” suggests a deliberately slow pace, which flies in the face of “defying the steepness.” And where does “gravel” come from? Sure, gravel bikes are all the rage right now, but El Toyonal is a paved road. Meanwhile, a human on a bicycle cannot be said to “soar,” and ChatGPT just tacked on the concepts of caprice and radical freedom without integrating them into the poem. The A.I. gives no indication (or I should say simulation) of even knowing what these terms mean.

It’s odd that this poem actually makes less sense than ChatGPT’s sonnet … it’s almost as though the chatbot blew all its computing cycles fighting with the meter. This poem is only a bit better than what GPT-3.5 had come up with, and undermines the sense that GPT-4 actually understands the structure of language. Maybe ChatGPT’s progress with sonnets is just due to imitation; after all, there’s vastly more training data available for that form.

(If you’re interested on comparing ChatGPT’s poem above to my own dactylic trimeter poetry, click here and/or here.)

ChatGPT art

I’ve never before delved into the artistic capabilities of ChatGPT, so I don’t have any benchmark by which to evaluate its progress over earlier versions, but you gotta start somewhere, right? As it happens, I visited my older (fledged) daughter recently and, following an incident involving a hot tub, she started messing around with ChatGPT and asked it, “Can you create an image of a tall skinny white man feeling faint after leaving a hot tub?” Here’s what it came up with:


When my daughter showed me this, I immediately pointed out that, perhaps based on some automatic effort to make the man good-looking, ChatGPT gave him too much upper body musculature to really be called “skinny.” I think “hunky” would be a more appropriate description. 

My daughter told ChatGPT, “Make him even skinnier.” Almost as if being sassy, the chatbot produced this:


My daughter prompted ChatGPT to try again without going overboard, and its next effort looks a lot like cheating:


Not only is this a copout, but the picture suggests an implausible scenario. If this guy felt faint after leaving the hot tub, and then took the time to go find a robe and yet still feels faint, why isn’t he either wisely sitting down, or sprawled out on the deck having passed out? Also note that part of his robe’s belt is missing.

My daughter went back to the original picture and told ChatGPT, “Make him skinny like a cyclist not like he is anorexic.” Here’s its response:


The cycling shorts are a cute touch, but not very realistic when you think about it. What cyclist wears his cycling shorts in the spa? And who said this guy just finished a ride? It’s not like cyclists wear their cycling clothes all the time. This hot tub could be at the guy’s home, or at a hotel he didn’t even bring his bike to. Meanwhile, the picture still fails to capture the physique of a typical cyclist … very few of the riders I know have pecs or biceps that big.

Moving along from the hot tub pictures, last week I didn’t have any cover art for my blog post, so (inspired by my daughter’s experiments) I decided to see what ChatGPT could come up with. I asked it to create a picture, in the style of William Pène du Bois, of a teenage girl using ChatGPT on a tablet. The result is a far cry from du Bois, and though I used it anyway, I received some constructive criticism from a reader that the picture was perhaps not quite appropriate for the top of my post. Thus, I replaced it (eight days after I had originally posted it) with a different one (more on this later ... see the Epilogue at the bottom of this post). Here is the original picture that ran at the top of last week’s post:


The issue with that picture is the girl’s bare shoulder ... a bit racy especially given her age. I didn’t really like that from the beginning. I asked ChatGPT to fix that, and make the girl’s cheeks less rosy, and make the cat more realistic, and it produced this:


I don’t know about you, but I find this second effort deeply unsettling. Her cheeks are just as rosy as in the first picture; her eyes look like an exaggerated attempt to appear as Western and doe-like as possible; and overall there’s just this air of uncanny-valley old-timey weirdness like you  get with the American Girl dolls. The picture is more like what Thomas Kinkade would create than Pène du Bois.

I asked ChatGPT to go back to the first drawing and try again without the bare shoulder, but to keep the clothing modern, and here’s what I got:


This isn’t so bad, but how is that clothing modern? Who wears overalls anymore, and big puffy, flouncy sleeves? The girl’s entire house looks antique. But my main issue is the weird non-words on the tablet display: “Ceenly crerrity” and “Ininty ccnvity” which bring to mind the strange strings of non-words that bots sometimes include in bogus comments on my blog posts. I find them unnerving.

To create new cover art for today’s post, I decided to scrap the Pène du Bois picture and start from scratch. I asked ChatGPT, “Please create a picture, in the style of Shawn Martinbrough, of a tall, blond, lean, middle-aged man, with a cat on his lap, wearing nice bluejeans and a black hoodie, using a tablet to visit the website www.albertnet.us.” What it came up with was almost exactly like what you saw at the top of this post:


Not a bad picture, though the guy isn’t exactly middle-aged, and the cat isn’t exactly on his lap. Also, the strange text at upper right doesn’t make much sense. I told ChatGPT, “That looks great, but please make ‘albertnet’ all lowercase and don’t show ‘SHAWN MARTINBROUGH’ on the tablet display.” This would seem to be a trivial tweak for ChatGPT, but to my surprise it produced a substantially different picture:


As I’m sure you’ll agree, this dude looks even younger than the original. And what’s with the www.ShMartinbrough and other textual debris? How is a misspelled URL helping anything? I responded, “Great, but please remove the www.ShMartinbrough from the picture.” Just a tiny change. It came back with this:


So now we’re back to the original guy (which is fine, I won’t miss that Vanilla-Ice-looking dude), but ChatGPT forgot I wanted albertnet in all lowercase, and what’s worse, the guy is now holding the tablet backwards, staring at the back of it! WTF? That’s like the old trope of the guy pretending to read a book but (unbeknownst to him) it’s upside down. I replied, “That’s good but it looks like he is holding the tablet backwards. Also, he is too young, I asked for a middle-aged man. Can you fix those things?” ChatGPT tried again:


It just gets worse and worse! The tablet is still backward, and now we have this bizarre shadow image of the guy. Did I ask ChatGPT to get all artsy-fartsy on me? I did not. And okay, the guy looks a tiny bit older, maybe a year or two tops, but it’s mainly the cat that looks old.  I complained, and the chatbot (which to its credit never gets irritated or frustrated) gamely regenerated and came up with this:


Sure, the tablet is fixed, but now the cat just looks angry, the guy has just as much hair, and check out that cleavage! He’s got, like, pneumatic pecs! He needs a fricking brassiere! And look at the size of those shoulders—he’s a monster! I told ChatGPT, “Can you tone down his pectoral muscles and give him a bit of a receding hairline such as many middle-aged men have?” Here’s its response:


Oh. My. God. It’s all gone downhill. His hairline is essentially unchanged; the unwanted text is back (this time morphing the two names to produce the nonsensical “Albertinbrough”); the dude’s pecs are just as huge as before; now his shirt is this kind of flimsy satin that’s practically lewd; his biceps are about to burst; and check it out, this brute is actually sitting right on his poor cat! How is the poor beast’s spine not crushed? And yet the cat seems perfectly stoic about the situation. Not very realistic. In A.I. terms this is a “hallucination” and shows how ChatGPT is still unable to sanity-check its creations. What’s shocking to us doesn’t seem wrong to the A.I. Do I need to specify that I don’t want the cat’s head to be bursting out of the guy’s groin?

I tried three more times to fix the picture, emphasizing a non-crushed cat, thinning hair, a man at least fifty years old, the build of a cyclist, and albertnet in all lowercase. While I was at it I asked to make the cat a tabby. ChatGPT kept trying, swinging wild at this point, ignoring first this instruction and then that, producing all manner of artwork but without ever meeting all of my simple directives:



For each picture, ChatGPT provided a caption telling a nice lie about the revision. For example, below the last picture it wrote, “Here is the updated illustration with ‘albertnet’ in lowercase, the man having the lean build of a cyclist, and a tabby cat resting on his lap. Let me know if there are any other changes you’d like!” True, the picture was updated, and that is a tabby, but everything else about this description is incorrect. So I went back to the very first picture and, using a different A.I. tool, manually scrubbed off the errant text so I could have something usable for the cover picture. ChatGPT, instead of a precision tool, had behaved more like a dartboard. And I suck at darts.

As with the poetry, ChatGPT seems to want to be the whiz-kid who can crank out something passable in almost no time at all, vs. thinking deeply and producing something that’s spot-on. ChatGPT’s fail-fast, iterative technique strikes me as almost the opposite of art. For blog post cover pictures, I’d rather commission my younger daughter to take a little time and create something of real value (as she has done for previous posts like this one, this one, and this one). She works much more slowly than ChatGPT, and isn’t at my beck and call, but I think the end result is far superior. I couldn’t get cover art for this post because she’s away at college and it’s dead week, but to compare her work to ChatGPT’s, let’s compare an earlier effort of hers, drawn when (at age seventeen) she hadn’t yet taken any college art courses:


I asked GPT-4 to create a black and white drawing of a hand holding a mechanical pencil and here’s what it came up with:


Should I need to remind the chatbot how many fingers a human has? And what’s with all the stray dots … are they fountain pen ink spills, or beads of black sweat flung from the brow of a six-fingered space alien? Tell you what, I’m sticking with human artists for now. They’re worth the wait.

Conclusion

Looking back at these last two posts, I would say the current buzz around A.I. is well warranted, given a) how quickly the technology is improving, and b) the ramifications—not all positive—of how we get information from the Internet and what we get when we task A.I. with creating what will pass for our own creative output. I guess I shouldn’t be surprised to see Gen-Z people using ChatGPT and even Microsoft 365 Copilot as routinely we’ve all been using Google all these years. Myself, I prefer old-fashioned web search tools because my answers will be more complete, more interesting, and may take me down interesting rabbit holes that (so far) I still have the patience for. As for creating prose, poetry, and art, ChatGPT strikes me as a powerful tool, but one we’d better be careful to reign in. A.I. still seems to put speed and convenience ahead of quality and reliability. My take-away: power to the humans! Stay ahead of A.I.!

Epilogue

Getting back to that kind of odd picture from last week’s post, I decided today to replace it. From the beginning I hadn’t liked how the girl’s shoulder was bare and her bra strap showing, and a reader complained about this along with the fact that this youngish girl seemed to be wearing a lot of makeup. I decided to also abandon the part of my prompt that said to employ the style of William Pène du Bois ... that just wasn’t working out. So this time I promited ChatGPT, “Please create a picture, in the style of Chris Riddell, of a 19-year old girl in modest, modern attire in a modern setting using a tablet.” To my suprise, ChatGPT refused, saying my request ran afoul of “DALL·E’s content policy.” I asked for details, which helped narrow it down to the style component of my request, and ChatGPT told me, “This might be due to ... closely emulating the style of a living artist like Chris Riddell.” This puzzled me, since Shawn Martinbrough (whose style ChatGPT happily emulated two days ago) is also living. So as an experiment I asked ChatGPT, “Please create a picture, in the style of Shawn Martinbrough, of a 19-year old girl in modest, modern attire in a modern setting using a tablet.” Here’s what I got:

Does that weird sweater, with the oversized collar, look familiar? It’s the same garment the very first drawing featured, of the teenager done in the Pène du Bois style! What part of “modest attire” is this chatbot not getting? I asked it, “Can you please try again but not have her shoulder exposed?” It generated this:

The caption ChatGPT gave the above picture was, “Here is the updated illustration, ensuring her shoulders are fully covered and her attire remains modest and modern.” False! I see a shoulder, a bra, and cleavage! I replied, “I can still see her shoulder and the strap of her bra. Can you fix that by giving her a garment that covers both shoulders and doesn't show any strap?” It gave me this:

Curses! Foiled again! And ChatGPT will only generate three pictures a day for non-paying users like myself, so I decided to call it a day and used the above picture atop last week’s post. It looks like we may need to wait until GPT-4.5 or GPT-5 for the amazing new technology involving pictures of women that don’t show a bare shoulder and a bra strap. Perhaps hundreds of developers are working on that problem even as I type this. Until that breakthrough is made, I will maintain steadfastly that ChatGPT cannot be held to possess intelligence.

Other albertnet posts on A.I.

—~—~—~—~—~—~—~—~—
Email me here. For a complete index of albertnet posts, click here.