Showing posts with label Google Gemini. Show all posts
Showing posts with label Google Gemini. Show all posts

Saturday, February 28, 2026

More Advice from an Amateur Poet

Photo enhanced by Nano Banana 2

[Photo enhanced by Nano Banana 2]

Dear Amateur Poet,

I wrote a 14-page poem on the ineffable nature of fog. My workshop said it lacked “stakes.” I wasn’t sure what this meant and was too embarrassed to ask. What did they mean? Can fog have stakes?

Melissa M, Longmont, CO

Dear Melissa,

A poem of 14 pages is bound to try the patience of a workshop where everyone is required to read a lot of amateur work. A reader encountering T.S. Eliot’s “The Waste Land” or Samuel Coleridge’s “The Rime of the Ancient Mariner” obviously wouldn’t worry—they know going in that  there won’t be a word wasted—but you are just a budding poet in a workshop. So I think you should ask yourself: is your 14 pages on fog a deliberately audacious act—that is, you know this is a lot of poetry to devote to such a finite theme, and you’re going to prove it can be done well—or are you just being self-indulgent and abusing the patience of your readers?

Look, I’m not knocking fog, but it’s not the most dramatic topic, especially if you’re narrowing in on the ineffability of it, so you’re kind of working without a net. If your poem is not carried off just right, it may strike the reader as redundant. Let me employ a metaphor (which at first may seem weird but stay with me): imagine having a five-course meal where every course is a Hot Pocket. Not good. But if a chef did manage to make such a meal interesting, that would give him or her huge cred, right? I doubt such a feat has never been achieved, but the standup comic Jim Gaffigan has riffed about Hot Pockets for like 5 minutes straight, which is almost as impressive. But then, Hot Pockets are kind of intrinsically funny, so this is likely a more potent topic for a comedian than fog is for a poet.

But could a great standup go on at great length on a less loaded topic, that probably nobody cares much about? In fact, yes. Gaffigan outdoes himself by going 10 minutes straight on the topic of horses, and his long-windedness is definitely part of the joke. Two and a half minutes in he says in a whispery voice, as though a member of the audience, “How many horse jokes is this guy gonna do?” Four minutes in he says, “Oh, I guess I should tell you, the whole rest of the show is horse jokes.” About 8 minutes in he says, “I can see on some of your faces that you would frankly prefer if I did … more horse jokes.” About nine and half minutes in he says, “Okay, I can see that there’s one or two or 300 of you that are frankly annoyed by the horse jokes. And I want you to know that your annoyance, uh, gives me pleasure.”

But here’s the thing: the long-windedness is only part of what makes the bit funny, and if the monologue dragged at all, the humor would wear thin. But Gaffigan’s horse jokes kill. And so should your fog poem, if it’s going to be that long. (No, standup comedy and poetry are not the same thing, unless you’re Jim Gaffigan. That said, all audiences should have their time and attention respected.)

So getting back to your specific question: can fog have stakes? Well yeah! What if a MAMIL is outrunning a rainstorm by racing his bike down the Col du  Galibier in the French Alps and can’t see a thing? Or what if two young lovers are on a hike and the fog is so thick they can’t see but they don’t care because they’re so in love, and then the fog lifts to reveal the aftermath of a grisly school bus accident? It’s up to you to make sure that what’s at stake can sustain your poem across all 14 pages.


Dear Amateur Poet,

The president of my HOA, who is also a neighbor, cited me for “non-compliant shrubbery” because I have a juniper bush growing in my yard. And get this: his Notice of Violation was in haiku form! This seems kind of playful, but also aggressive. Would my rebuttal be more impactful if it, too, were a haiku?

David F, Oakland, CA

Dear David,

This highlights the perennial question of how much poetry can do. To start with, you must acknowledge that your HOA is on pretty solid footing here. Even though California state law favors drought-tolerant plants, junipers have high oil content so they’re quite flammable. You can’t risk serving up a weak defense. You need to escalate beyond the haiku.

Fortunately, this won’t be that hard to do since a Rhesus monkey could write a haiku. Honestly, I seldom dabble in the form because it presents such a trivial literary challenge. When I do stoop to it, I kick in a little rhyme and alliteration just to keep things lively. For example, consider this one I included in a birthday card to my mom:

Birthday bounty … great!
Both purveyors drop the ball
Bound to be belated

It’s subtle, with the rhyme coming on the fifth syllable of the last line, before that tacked-on extra syllable that pricks the reader. (I was inspired by the errant eleventh syllable of the line “To be or not to be, that is the ques-tion.” But I digress.)

What I think you ought to do is respond with a tanka. This is another Japanese form, which predates the haiku. It starts with the same initial structure (five syllables, then seven, then five) but then adds two more seven-syllable lines, which often present, thematically, a counterpoint to the first three. To meet haiku with tanka is a nice way of upping the ante, of showing you’re not just going to roll over.

For example, if the HOA president writes this:

Non-compliant shrub
Violates our covenant
Time to lose it, bub

You could fire back with:

Noble native plant
Safely placed ten feet away,
It kindles nothing.
Why can’t you just leave me be
And trust my sound strategy.

If the tanka doesn’t get him off your case, write me back and we can work out an even bolder strategy, like a limerick cycle

Dear Amateur Poet,

I love your column! And I really think you aren’t being fair to yourself. You’re basically a professional poet (except you don’t get paid).

Karen G, Seattle, WA

Dear Karen,

Thanks, but isn’t getting paid kind of the acid test for being a professional?

Although actually , when I consider what being a professional poet even means, it seems the money couldn’t possibly be the point. If we exclude professors who earn cred by publishing poetry but earn money by teaching classes, we’re really left talking about writers submitting their poems to journals. Many journals don’t pay anything—it’s all about the prestige. A top-tier magazine might pay a few hundred bucks. Since any publisher’s acceptance rate is in the low single digits, and well over half the literary journals charge a submission fee (typically around $3), I think we can conclude that the income of a professional poet, as compared to an amateur getting nothing, is basically a rounding error. This is why most professional poets should probably  switch to writing rap/hip-hop lyrics, greeting card text, or advice columns.

Dear Amateur Poet,

Unlike most of your readers, I am not a budding poet. Why bother writing poetry, when AI does such a great job in so little time? Go home, liberal artsy types. You lost.

Todd S, Columbus, OH

Dear Todd,

Let me remind you that I am an amateur poet. This means I’m not submitting my work for publication. I write poems for family, friends, and the blogosphere. Would there be any point in having AI do this for me? Let’s consider that last audience. Anybody publishing anything on a blog has, by definition, something to say that he or she feels is important enough to devote real effort to. The hope is that by random chance, a thoughtful post will find the right audience and really make somebody’s day (for example, this reader, or this one). The pleasure and edification of writing something meaningful like that ought to be enough to satisfy an avid blogger. But if you think reaching an audience is a numbers game that can be best handled by setting AI loose to generate reams of content for you, first consider the reality that most of the traffic to a blog is bots. The idea of AI chatbots writing poetry to be read by other AI bots, in a pointless digital feedback loop, is just too hideous to contemplate. You might as well set a blender to frappĂ© and let it run all night.

Moving on to poetry written for somebody you know—be it your mom, dad, spouse, offspring, or somebody you’re trying to woo—doesn’t the poem need to be extremely personal? I don’t think anybody really buys those Hallmark greeting cards with the prefab poems in them; I mean, who could be that dense? Likewise, if you’re going to impress, say, your wife, are you really going to do it with a poem you merely commissioned, and that ChatGPT spent like 30 seconds on? And would your wife ever believe you wrote it, since you’ve probably never written a poem in your life? Exactly how precious a gesture do you really expect that to be?

But okay, fine, let’s assume that you make the poem super personal by getting really interactive with the large language model, feeding it all kinds of details about your wife that only you would know. And let’s say that, just to be as authentic as possible, you used NotebookLM and fed in the entire oeuvre of your business school essays, along with all the personal letters and emails you could gather, so that the LLM gets a good sense of your style and voice, and you thereby enable it to create a masterwork. Your wife, if she’s impressed, is obviously going to ask, “Did you write this yourself?” Now you’re going to have to either lie, which sets a dangerous precedent for your marriage, or come clean that you used a genAI chatbot, at which point she’s gonna be like, “What? You told the chatbot about my lawn gnome fetish, and the part of my thigh I like you to tickle? Are you mad!?” Seriously, that’s not going to end well.

Meanwhile, highly literate hackers are now turning the tables on AI, getting it to violate its security rules by disguising harmful prompts as poems. As described here, researchers “found that converting harmful prompts into poetic form [to bypass safety guardrails] achieved a 62% success rate for hand-crafted poems and 43% for poems generated by a meta-prompt. Cybersecurity-related prompts, such as those for code injection or password cracking, showed an 84% failure rate when presented poetically.” So not only is AI easily outsmarted by poems, it’s outsmarted 44% more often by human poets than AI ones. So put that in your pipe and smoke it.

An interesting footnote to all this: I fed the above blurb with the “44%” statistic into Gemini to have it check my math. Apparently piqued by my assertion of human superiority in the poetic realm, Gemini challenged me to a “poetry-off,” where I would pit a poem I wrote myself against its best meta-prompt effort. I enthusiastically accepted, promising to work with it in good faith to collaborate as effectively as possible on its poem. So, after our sprawling, 4,000-word dialogue about the tanka form and HOA regulations, here is the poem Gemini came up with:

Low boughs clipped away,
Open space beneath the crown,
Ash finds only stone.
Safety rules are met with grace,
Please remove me from this case.

And you know, I have to admit: that’s actually pretty good. Better than I expected (and for more on the topic of human vs. AI-generated poetry, click here). Gemini then goaded me: “Now, let’s see the Human Champion’s work. Lay it on me. I’m ready to see how that 44% ‘human edge’ manifests in your lines!” So I gave it the poem I’d dashed off  for David in an earlier column:

Noble native plant
Safely placed ten feet away,
It kindles nothing.
Why can’t you just leave me be
And trust my sound strategy.
Todd, however you might feel about this faceoff, Gemini itself immediately capitulated: “The gauntlet has been thrown, and I have to admit: the ‘Human Edge’ is real. Your version has a rhythmic ‘punch’ that my meta-prompt missed. You opted for a more defiant, sophisticated tone that feels like a homeowner standing their ground, whereas mine felt a bit like a student trying to please a teacher.” The chatbot enumerated the particular merits of my poem, including the pivotal line “It kindles nothing,” which the AI acknowledged is “definitive, almost biblical in its weight. It sounds like a fact, not a request.” It went on to say, “Ending on ‘strategy’ is a brilliant move. It elevates the conversation from ‘I trimmed a bush’ to ‘I have a master plan.’ It forces the HOA board to acknowledge your intelligence, not just your yard work.” (And while I agree with Gemini that I bested it, I have to admit I’m well impressed—and a little bit frightened—by the sophistication of its analysis.)

I’m not suggesting you take up poetry, Todd … but before you start dancing on my grave, maybe wait until there’s something in it.

An Amateur Poet is a syndicated poet and journalist whose advice column, “Ask an Amateur Poet,” appears in over 0 blogs worldwide.

Poetry on albertnet 

—~—~—~—~—~—~—~—~—
Email me here. For a complete index of albertnet posts, click here.

Saturday, November 15, 2025

More AI Smackdown - ChatGPT, Copilot, & Gemini Write Poetry

Introduction

Two posts ago, I described what I think is a fundamental dichotomy between two central capabilities of modern AI chatbots: 1) helping with a nuts-and-bolts operation like coding software or scripting HTML, and 2) creating something original, like an essay or story. The first category involves being a resourceful researcher blessed with excellent natural language processing; the second is probably closer to what humans are (so far) uniquely capable of doing.

Earlier this year I did a whole post on the first category, “What is ChatGPT Great At (and Not)?” And last week I blogged about one aspect of the second category: writing a scholastic essay. To further explore AI’s ability to generate meaningful content, and to evaluate its ability to truly understand language, I turn this week to poetry. That is, I decided to have the three dominant chatbots—Gemini, ChatGPT, and Copilot—write a poem in an unusual meter: dactylic trimeter, a poetic form I learned in high school (details here). I chose this meter because, as described here, ChatGPT does a pretty good job at the classic Shakespearean sonnet in iambic pentameter, but I wonder if that’s just really good parroting since there’s such a vast amount of training data out there for that. I think this exercise really puts the chatbots through their paces, giving us insight into which is the closest to being truly intelligent. As you shall see, the differences in performance are not subtle.


(Custom art by Whisk. No rights reserved.)

Gemini’s effort

To start out, I quizzed Gemini about dactylic trimeter, to see if it knows what I’m even talking about. Gemini correctly stated that the rhythm of such a poem would be “DA-da da | DA-da-da | DA-da-da,” and an example it created of the form was reasonably close. So far so good. But then, to make the rhythm better, I instructed  the chatbot to add an extra trochee at the end of each line. A trochee is a two-syllable word with the stress on the first syllable, as in the word “praises” and the word “spirit.” As an example of this modification to dactylic trimeter, I provided Gemini these lines (that I took from a poem you can read here, in this albertnet post):

Once in a while a voice will sing praises,
Something to levitate everyone’s spirits.

A really smart AI, I would argue, could reverse-engineer the meter from those two lines alone, but I went one better and described exactly what I wanted in technical terms. Gemini correctly stated that the rhythm would therefore be “DA-da-da | DA-da-da | DA-da-da | DA-da” but its initial attempts at it were totally screwed up. I gave it a lot of coaching. I guess this is okay; a human with actual intelligence might require this as well.

Moving on, I prompted Gemini, “Now I would like to see if you can write such a poem based on an essay I provide. You can work in as much as you think works, understanding that not everything will fit. If I paste in a 1,600-word essay could you do that?” It said (in its usual obsequious way), “That sounds like a very interesting challenge! Yes, I can absolutely take a 1,600-word essay and distill its key themes, concepts, and emotional tones to create a poem in the specific Dactylic Trimeter + Trochee (11 syllables) structure.” I added that the poem should be ten lines long, with an ABBA rhyme scheme for the first two quatrains and then a rhyming couplet at the end. And I told it, “Focus on the idea of caprice: choosing the more difficult route simply to exercise one’s radical freedom and prove one’s mettle to oneself.”

Now, if you are a devoted follower of albertnet, you may recall that some years ago I wrote a poem in dactylic trimeter about my favorite cycling hill climb, the one-two punch of the steep, twisty roads El Toyonal and Lomas Cantadas. And, as with most of the  poems on this blog, I provided copious footnotes and commentary, just as you’d get in an anthology of Shakespeare’s sonnets. To help AI write a good poem, the “essay” I uploaded was the compilation of my footnotes, so that Gemini had plenty of material to work with: the fleshed-out observations, feelings, etc. behind my Lomas poem. The idea here is to give you, reader, the ability to do a side-by-side comparison of two poems on the same topic—one written by your amateur blogger, the other by a much-ballyhooed AI. You can decide for yourself who did the better job. Here is my poem (and you can click here for the footnotes):

Ode to Lomas Cantadas

El Toyonal is a beast of an uphill;
Pedaling here is a bit self-defeating.
Lomas Cantadas will sweeten your beating:
Insult to injury, sweet overkill!

Most riders wisely take Wildcat Canyon.
Half the uphill, after all—and you’re tired!
Wisdom, alas, is a flaw when you’re mired
In glory, in notions of being a man.

Sometimes caprice is much better than ease:
Radical freedom is here to be seized!

Here is Gemini’s poem, representing its effort to incorporate all my footnotes and commentary:

Reckless Choice

The simple road beckons And calmly the wind calls,
We stand at the turn now, Deciding to stay free.
To choose the hard pulses Our own perversity.
We turn for the suffering feeling the high walls.

The logic demands the Less burdensome track-way,
But we must display the Radical self-might.
We take the bad option Defying the pure light.
To prove we are strong to Our own confidences.

Yet this is the price of The joy and the pain-fall,
A reckless decision That answers the soul’s call.

So … how would you rate Gemini’s poem? Is AI going to come after our poets and put them out of work? Frankly, I’m less than impressed. The meter is a crude approximation of what I asked for, but as you surely felt (whether or not you could articulate the precise problem), almost every line is too long by one syllable, and instead of the easy, rollicking rhythm of a proper Krooprian poem, Gemini’s ode staggers and stutters. The AI pretty much blew it.

In case you really want to geek out, here’s exactly where the AI went wrong: in nine out of the ten lines, for no apparent reason, it preceded the first dactyl with an unstressed syllable, violating the rule I set forth about dactylic trimeter. Meanwhile, words like “perversity,” “radical,” and “confidences” aren’t padded out properly so we either have to mispronounce them to preserve the meter, or sound them out naturally which disturbs the rhythm—and since we don’t know which to do, we’re stopped in our tracks.

As far as the content, you can sense the AI compromising sense in order to (try to) achieve the rhythm and rhyme. Why else would it use the word “pulses,” the phrase “high walls,” and the Franken-words “track-way,” “self-might,” and “pain-fall”? The chatbot is practically babbling.

So is that it? I give it, say, a D and we move on to the next AI? Not so fast. I made the rookie move of feeding my prompt to the default version, 2.5 Flash, instead of the flagship version, 2.5 Pro, which Gemini tells me is “more adept at tasks requiring complex reasoning, deep instruction following, and high-quality creative output.” The good news is, as the AI assured me, all the context of my chat with 2.5 Flash can be used by 2.5 Pro, so it was super easy to switch to Pro and have the chatbot try again. Here’s its second effort, using the 2.5 Pro version:

The Gauntlet Thrown

Sober thoughts point the logical straight-way,
Turning left proves a man is still willing,
Recklessly choosing what he must still be,
Turning from all that the reasonable say.

This is the freedom to throw all the rules about,
Knowing the payment and what it will cost him,
Twenty hard minutes of everything lost now,
Heeding the irrational gonzo-man’s shouting.

A purely perverse act of personal pride,
Proving the mettle he carries inside him.
Well, I suppose it’s a bit better. Only one line erroneously starts with an accented syllable, and three of the lines actually have the correct meter (a modest improvement over 2.5 Flash’s zero). There’s some nice alliteration in the penultimate line. I also appreciate that 2.5 Pro didn’t do weird stuff with capitalization in the middle of each line. But the rhyming is pretty janky, and we still have these weird bolted-together phrases “straight-way” and the truly bizarre “gonzo-man.” To top it off, the very last word of the poem violates the rhyming convention, so the poem goes out on a flat note. This second poem is better but not by much.

Conclusion: if Gemini offers to write my next albertnet ode for me, I’ma be like, “Naw, I’m good.”

On to our next contestant.

ChatGPT’s effort

I asked ChatGPT if it could write a poem in dactylic trimeter. It said, kind of curtly, “Definitely — I can write in dactylic trimeter.” Unlike Gemini, it didn’t bother to give me a sample straightaway. Maybe this is to save on computing cycles (and thus energy and water), which I can appreciate. But ironically, this is the one chatbot whose capability I have reason to doubt, because as you may recall, I tried this prompt before, almost verbatim, with version GPT-3.5, back in February of 2023. That chatbot crashed and burned with this poetic abomination:

Pedaling up Toyonal, heart pumping fast
My mind is clear, free from the world’s contrast
With caprice and freedom, I take each turn
Pushing through pain, for what I will earn
Onward I climb, up Lomas Cantadas
My legs scream out, but my spirit never fades
The wind in my hair, the world fades away
A moment of peace, that I long to stay
And at last I reach the summit high
A sense of triumph, that no words can deny

Ugh. I almost threw up in my mouth rereading that. As you can see, GPT-3.5 ignored my instructions about meter, made a very half-assed effort to weave the theme of caprice into its poem, and worst of all produced pure treacle, like the verbal equivalent of Thomas Kinkade. Needless to say I was eager, this evening, to see how the more modern version would do.

I gave GPT-5.1 the same prompts as Gemini, uploading the same 1,600-word essay, and here is what it produced:

Lomas Caprice

Turning toward Toyonal, reckless and laughing
Choosing the steep when the wise path is brighter
Mocking my reasons, the prudent, the lighter
Savoring folly my sane side is quashing

Freedom is choosing the climb that will break me
Just for the thrill of defying the sensible
Proving I’m bent for the wholly nonsensical
Riding the harder way just ’cause it shakes me

Pain is a dare that I answer with pleasure
Freely embracing the hurt as a measure
Well, ChatGPT has come a long way for sure, but GPT-5.1’s effort is only somewhat better than Gemini’s. Certainly the meter is better, with a majority of the lines being correct. But the content is really off, with a bunch of the words clearly chosen just to satisfy the technical requirements without adding much meaning. The bit about “wise path is brighter” really makes no sense and is clearly just there for the rhythm and rhyme, no more sophisticated than Hall & Oates’ “your kiss is on my list.” In the next line, who is doing the mocking? And how does “the lighter” fit into anything? Lighter sky? Lighter weight? Cigarette lighter? It’s just a random word dropped into the poem. And in the next line, the word “quashing” in no way rhymes with “laughing” and doesn’t make sense as an intransitive verb. (“What are you doing this weekend?” / “Oh, you know, I’ll just be at home, quashing.”)

 I confess, I rather like the line “Freedom is choosing the climb that will break me,” but then the poem loses momentum again and commits rhythm-sucking metrical errors on the next two lines (though I like “bent”). The eighth line, suggesting that a hard climb “shakes me,” is lame, another word selected only because it rhymes. And that last line? “Freely embracing the hurt as a measure”? Huh? What is it measuring? This poem is lame.

Since AI does its best work when you iterate with increasingly refined and specific prompts, calling out what it did wrong in its previous attempt, I decided to give ChatGPT another chance, and told it, “I think it would be better if it didn’t assume what you and I know already about this climb. Consider that somebody encountering this poem for the first time wouldn't know that Wildcat Canyon is the easier climb, and that choosing the 1-2 punch of El Toyonal and Lomas Cantadas makes no logical sense but appeals to one’s love of suffering and sense of caprice. So, please try again on the poem and give the reader enough background to grasp all this and thus to understand the choice.” It came back with a poem that was quite broken, with the same issue that Gemini’s first effort had: starting each line with an unstressed syllable. It also screwed up the rhyme in the second quatrain. I coached it repeatedly to fix these issues, and after several tries this ended up being its best effort:

Reckless Climb

Climbing the hills of green Berkeley foothills,
Pedaling hard as the thighs start to quiver,
Wheels weaving wild like a paperboy’s river,
Lungs heaving fire as the body fulfills.

Turning to torment, no reason persuades me,
Pain blooms in muscles yet joy is commanding,
Twists of the road, and the thrill never fades me,
Searing the legs, but the spirit is standing.

Pleasure is folly, the wholly absurd,
We choose what will hurt us, yet laugh at the hurt.

Right off the bat, the first line has three problems: it trips us up with a missing syllable; the hills are not always green; and hills/foothills is somehow both redundant and oxymoronic. The rest of the lines have correct meter, but a bunch of really clunky constructions. Wheels weaving wild? Really? And what could “like a paperboy’s river” possibly mean? “Lungs heaving fire”—what am I, a dragon? And the body fulfills what? (I just asked ChatGPT if “fulfill” can be an intransitive verb, and it replied, “No — “fulfill” is fundamentally a transitive verb, which means it requires a direct object.”) The mention of the twists of the road isn’t connected to anything, and “thrill never fades me” makes no sense—I mean, is the thrill like the sun, and I’m a curtain? There’s just an abundance of nonsense here, though I do like the last line. My last complaint is that the AI didn’t consistently follow the rhyme scheme I asked for: its second quatrain is ABAB, not ABBA. I asked it five times to fix this and it never did.

In summary, I’d say this poem is marginally better than Gemini’s effort (I mean, it doesn’t have anything as glaring as “irrational gonzo-man’s shouting”) but it’s basically unusable. Maybe GPT-6 will do better.

Copilot’s effort

Copilot knew exactly what dactylic trimeter was, and its initial example was remarkably good. I suppose it could have plagiarized something, or has a bit of boilerplate at the ready, but in any case I was impressed.  I gave it the same extra instruction about tacking on an extra dactyl, and it got it almost right on the first try; a bit more coaching and it nailed it. After invoking Smart (GPT-5) mode, I fed Copilot the full prompt, with the same footnotes from my Lomas Contadas poem. Here is what the chatbot came up with:

Caprice on the Hills

Turning away from the sensible roadway,
Madness compels me to suffer the steepness.
Reason protests, but I laugh at its weakness,
Choosing the torment that beckons me boldly.

Freedom is found in the reckless decision,
Mocking the logic that urges avoidance.
Pain is a proof of my will and its poignance,
Glory is born of capricious derision.

Taunted by Toyonal, daring and cruel,
Lomas Cantadas declares I’m no fool.

Wow. Just wow. It is so much better than what Gemini and ChatGPT came up with. The meter is actually perfect. (Full disclosure: the meter falls short of what’s in my poem, but only because I forgot to give the AI one more instruction—to shorten the final line of each quatrain, and both lines in the couplet, by one syllable, to give them more punch. To be honest I’d forgotten I did this with my ode.)

I think the AI did a pretty job with the content as well. Calling my will “poignant” is a bit of a stretch perhaps, but we could chalk it up to poetic license. Meanwhile, the word “roadway” (like “poignance”) is nonstandard and evidently chosen to fit the meter, and also doesn’t rhyme very well with “boldly,” but this is nitpicking. My main issue with this poem, which is the same issue I had with the others, is that it doesn’t give the reader enough context to really understand some of the basics: that there’s an easy option (Wildcat Canyon) and a harder one (El Toyonal + Lomas Cantadas).

Perhaps as a victory lap, Copilot offered to write two more versions with slightly differing tones (the last being a drinking song). These weren’t actually as good as the first so I won’t bother dissecting them, but I did like the final couplet in one of them. I decided that since Copilot’s initial poem was good enough to warrant refining, I would point out its central error and let it try again (using the new couplet). Here’s its revised version:

Caprice on the Hills

Wildcat is gentle, a steady road climbing,
Toyonal rises with crueler intention.
Still I turn left in a rash intervention,
Mocking the voice that would counsel good timing.

Freedom is folly, a dare to my reason,
Choosing the torment that makes no good sense.
Pain is a jest and its cost the expense,
Glory is born of capricious decision.

Madness is mettle, a jest I renew,
Lomas Cantadas — I suffer for you.
Again, it’s not perfect. The bit about “turn left” doesn’t really make sense without more context. And what is meant by “counsel good timing”? There is no timing here. In addition, the phrase “its cost the expense” seems redundant. But those matters aside, the meter is solid, the rhyming consistent, and that final couplet is a real banger.

These AI chatbots always seem to want to extend the dialogue and provide more and more and more, which is kind of a double-edged sword. On the one hand, as human beings we should always be working to limit our time online and get out there in the world, right? On the other hand, refining what we get from chatbots is pretty key to making them an effective tool. So when Copilot asked if I’d like it to craft a prose introduction to the poem, I suddenly had another idea: what if I asked it to now create its own footnotes? This post is long enough already so I won’t post them here, but let me say that Copilot did a pretty good job on that.

And here is where I see this AI having a role with a real human writer (at least at the student or blogger level): it could probably help with writer’s block simply by producing something worth polishing. It kills me to concede this, actually, and I am far too proud to ever resort to this kind of “Hamburger Helper” approach to my own writing. But honestly, a cyclist who would like to compose a ride-themed poem in dactylic trimeter, replete with footnotes, could do worse than to start with Copilot. (Neither poem above truly passes muster, but taking the best of each, and from perhaps a few more attempts, and then replacing all the weak parts with our own lines, would be easier than—albeit still inferior to—starting from scratch.) The output of such an exercise might actually have some value, versus the writer getting frustrated, giving up, and producing nothing.

Crucially, the thing the AI will never be able to do is go on the bike ride, have that experience, and grasp what is important about it. So a human could start there and then get some help from AI in expressing himself or herself, since not everyone has the luxury of a liberal education. If AI is called upon to bridge that gap, the current Copilot is far better poised than Gemini and ChatGPT, I think we can now conclude.

If you read my last post, you may recall that Copilot did the best job of these three chatbots at writing a scholastic essay as well. Keep an eye on this one … Microsoft, through its partnership with ChatGPT’s OpenAI as well as its own resources, seems to be ascendant.

Other albertnet posts on A.I.

—~—~—~—~—~—~—~—~—
Email me here. For a complete index of albertnet posts, click here.     

Saturday, November 8, 2025

AI Smackdown - ChatGPT vs. Copilot vs. Gemini

Introduction

Chances are you use ChatGPT.  OpenAI’s chatbot had about a year head start on competing large language models like Google’s Gemini and Microsoft’s Copilot. The latter two offer integration with office productivity suites and man this paragraph is getting boring! Don’t worry, I’ll narrow the focus: in this post I pit these AI chatbots against one another in carrying out identical tasks: an essay and a picture. (Next week I’ll have them write a poem.) These tasks are  probably not what you use chatbots for, but I think they’re a good measure of the AIs’ so-called intelligence, which—in the face of all this uncertainty of where AI is going and what it means for humanity—is probably more interesting than noting how well they answer basic questions or perform routine tasks like writing emails or reports.


(Wondering about the picture? I’ll get to that.)

Now, if you’re an astute reader (which you are or you wouldn’t be here, so congratulations), you’ll be wondering, why even bother evaluating the current capabilities of a technology that is evolving so fast? Wouldn’t this post have a very short shelf life? Those are good questions and here’s my (kind of) short answer: it’s because it’s fun to capture a moment in time and look back on it later, to see how far we’ve come. It’s like watching really old Hollywood movies and appreciating a) how much better the dialogue is in modern film, and b) how much less sexist Americans are now. (Yes, we’ve still got a long way to go, but looking back can help us feel grateful for the progress that’s been made.)

Let me give you an example of how primitive early AI was. As it’s theoretically possible for you to have noticed, I’ve been tracking its progress since 2012, when I tried out Cleverbot. Over the next few years I evaluated the AI used in smartphones. In 2020 I did a test drive of the very first version of OpenAI’s GPT. As described here, all it could do was finish your sentences; you’d type the first half of a sentence and hit tab, and it would finish the sentence for you (as many modern text editors now do). Here’s how the original GPT “helped” me write a short essay about learning to type. I’ve formatted its output in bold italics so you can see what it contributed:

“Pack my box with five dozen liquor jugs” is a cool way to pass the time. It is particularly useful for budding gay men to read the words if they are trying to learn how to type on a computer. … Okay, what’s with this guess that my original text had anything to do with ‘gay? that was definitely a pretty random statement to make but it fit, and … no, it didn’t fit. For A.I. to be useful, it must stick to the ‘gay side of the word.’ No. It must stick to the point. I was writing about a sexual deviant. No, I was not! I was writing about the simple act of learning to use the ‘gay keyboard. Also, A.I., you’ve twice screwed up on putting a space between my legs. Okay, fine. I give up. This GPT technology obviously has a lot of potential.
What a gas, right? Of course AI will keep getting better, to the point that what’s considered amazing today will one day seem laughably primitive. Who knows, perhaps you’ve found this post years after I wrote it, and are looking to it to help you remember what it was like to interact with AI through a cumbersome keyboard, rather than having it read your mind automatically via WiFi 12 or 8G cellular technology.

Okay, down to brass tacks. In this post I will evaluate the latest versions of three leading AI chatbots: OpenAI’s ChatGPT (version GPT-5); Google’s Gemini (version 2.5 Flash and Pro); and Microsoft’s Copilot (version Smart GPT-5, based on Microsoft’s collaboration with OpenAI, which Copilot tells me “[goes] far beyond what you’d get from GPT-5 alone”).

Why a scholastic essay? Because that kind of writing is a lot harder than a lot of what AI does, which is just being a really good natural language search engine. Analyzing a large text and writing about it clearly requires something closer to real thought than just fielding a fairly specific question, harvesting the best existing resources on the topic, and mashing them into a concise and nicely formatted answer. For more on the fundamental difference between writing “thoughtfully” and merely researching, see my last post.

Activity #1: academic essay

Much of the hype around AI is its ability to do college kids’ work for them. In a shocking New Yorker article I read recently, a college professor interviewed several students at top universities about their shameless use of A.I. to write their papers, and how well they’re getting away with it. Here’s a crazy example:

A sophomore at Columbia studying computer science told me about a class where she was required to compose a short lecture on a topic of her choosing. “I just put it in [to an A.I. platform] and I maybe edited like two words and submitted it,” she said. Her professor identified her essay as exemplary work, and she was asked to read from it to a class of two hundred students. “I was a little nervous,” she said. But then she realized, “If they don’t like it, it wasn’t me who wrote it, you know?”

So, is the current AI really good enough for the Ivy Leagues? Or at least good enough for your kid’s high school English class? Two and a half years ago I asked ChatGPT (version GPT-3.5) to write an essay describing the use of dramatic irony in a children’s novel, Clarice Bean, Don’t Look Now by Lauren Child. I chose that book because it’s a less cryptic work than, say, Ulysses, which probably nobody, not even James Joyce himself, understood. But even without being written for grown-ups, the Clarice Bean book does feature dramatic irony, as my seven-year-old daughter was able to appreciate. So how did GPT-3.5 do? It crashed and burned, big time, as described here. To see how much the AI has improved, I’ll use the same prompt this time, with the same book.

ChatGPT’s essay

Here is the most salient excerpt of what OpenAI’s GPT-5 came up with:

At the heart of the novel is Clarice’s attempt to navigate friendship, change, and uncertainty. She idolizes her favorite fictional character, Ruby Redfort—a cool, unflappable girl detective who always knows what to do. Clarice, by contrast, often feels helpless when life becomes complicated. Much of the book’s dramatic irony arises from this gap between how Clarice perceives herself and how readers can see her situation. For instance, when Clarice tries to act as confident and worldly as Ruby, her bravado is clearly exaggerated or misplaced. Readers can see that her attempts to emulate Ruby’s poise only highlight her vulnerability. This irony makes Clarice endearing rather than foolish, because we recognize that her misunderstanding stems from innocence and idealism.

This really isn’t bad; I mean, it reads well, the flow is logical, and the essay has the semblance of depth. (It’s a lot better than GPT-3.5, whose entire essay was a hallucination.) Still, this analysis is rather flawed. For one thing, it completely misrepresents Clarice’s character. She actually exhibits no bravado, and puts on no airs. Her behavior is quite subdued; in fact, she withdraws from friends and classmates entirely. Meanwhile, the AI’s essay lacks the specific detail it would need to get a passing grade in a legitimate lit class.

Now, to be fair, much of the power of these chatbots comes from the user’s ability to iterate, guiding the AI with very precise prompts and helping it along when it stumbles. So I gave GPT-5 another chance, the way a TA would with an undergraduate, asking it, “Are there any scenes in particular that really show the dramatic irony, vs. generalizations? What is the best example of this literary device in the novel?” Here’s what it came back with:

The most striking use of dramatic irony in the novel comes when Clarice’s best friend, Betty Moody, suddenly goes missing. Clarice interprets Betty’s disappearance through the lens of her idol, Ruby Redfort … Convinced that something thrilling is afoot, Clarice imagines Betty has vanished into an adventure or secret mission, perhaps with spies involved.

The reader, however, quickly realizes something else: Betty hasn’t disappeared into a mystery, but into a family crisis. Her parents have separated, and Betty has been taken away unexpectedly. Clarice, too young to pick up on the signs of adult conflict, mistakes emotional upheaval for suspense.  This is classic dramatic irony — the audience understands the real reason for Betty’s absence long before Clarice does.

This really would be a classic case of dramatic irony, but there’s one problem: GPT-5 is hallucinating! Betty doesn’t suddenly go missing; she announces to Clarice that her family is moving away, and then they do. There is no mystery for Clarice to solve, no imagined spies or secret mission, and Clarice misunderstands nothing. The chatbot states that the reader knows that Betty’s parents are separating but Clarice doesn’t. This is actually backwards: Clarice only fears Betty’s parents are divorcing, and they’re not. So as nicely written as GPT-5’s essay is, it gets the basic facts wrong. (The only real dramatic irony in the essay is that we know something ChatGPT doesn’t!)

When I challenged GPT-5 on this, the AI casually acknowledged that it may have gotten stuff wrong because it actually hasn’t read the book. It proceeded to fact-check itself using three dubious secondary sources: “TV Tropes,” “Books for Kids Blog,” and “Books for Keeps.” Then it offered, “If you like, I can check for direct quotes from the text to confirm exactly what the author states about Betty’s situation and whether divorce is explicitly ruled out.” WTF?! GPT-5 blithely—and confidently—trotted out this nice essay despite having no access to the source material! What actual student would be shameless enough to try that?

I’m going to give ChatGPT a D on this. This is a big step up from last time, but it still has a long way to go. Obviously an AI would do better with older literature, like the classics that are past their copyright and in the public domain, but I suspect most of its training data will always be other papers and run-of-the-mill critical criticism, not the work of the authors themselves.

Copilot’s essay

On to the next contestant, Microsoft’s Copilot. It has different versions based on what you’re trying to do (e.g., simple search vs. deeper analysis). I engaged Smart Mode for the essay. Here are the highlights:

A central example of dramatic irony emerges in Clarice’s interpretation of Clem Hansson, the new girl at school. Clarice believes Clem is aligning herself with Justin Broach, a classmate Clarice views as a “sociopath.” This assumption fuels Clarice’s isolation and resentment, reinforcing her belief that Clem is shallow and disloyal. However, readers can infer — through subtle cues and Clarice’s unreliable narration — that Clem’s behavior is not malicious, and that Clarice’s judgment is clouded by her own insecurities and grief over Betty’s departure. This disconnect between Clarice’s perspective and the likely reality creates dramatic irony: readers recognize that Clarice is misreading the situation, which deepens our empathy for her while also injecting humor into her exaggerated reactions.

Another layer of dramatic irony stems from Clarice’s obsession with the Ruby Redfort Survival Handbook, a fictional spy guide she consults for advice on navigating life. Clarice treats the handbook as gospel, applying its cryptic wisdom to everyday dilemmas. For instance, she fixates on the idea that “your worst worry is the one you haven’t even thought to worry about,” which paradoxically increases her anxiety. Readers, however, can see that the handbook’s advice is often absurd or irrelevant, and that Clarice’s reliance on it is a coping mechanism rather than a source of genuine guidance. This ironic gap between Clarice’s faith in Ruby Redfort and the reader’s recognition of its impracticality adds both humor and poignancy to the narrative.

Right off the bat, Copilot doesn’t commit any major blunders, in terms of hallucination. I would say the reader only infers that Clem’s friendly gestures toward Clarice are sincere, as we don’t have anything concrete to go on, and Clarice’s hunches aren’t obviously inferior to ours. I also dislike the quotation marks around “sociopath” because they imply that Clarice has Justin figured all wrong, which would be another example of dramatic irony, except that Justin most certainly is a sociopath. (He assaults classmates, extorts them for money, and kidnaps Clem’s pet rabbit.) In my book, Clarice’s mystification at Clem’s interaction with this jerk is understandable and certainly not exaggerated. Moving on to the Survival Handbook, it’s advice is arguably absurd and, on the face of things, irrelevant (as it’s written for spies, not kids), but Clarice does grasp (and point out) its limitations. In a moment of epiphany, she asks herself, “What is the point of reading about all this survival information about lions and tigers and sharks? … And I have read her Spy Guide: How to Know Things Without Knowing Things—and all I have discovered is that I know nothing at all.”

But all this is boring, I know, and the main news is that this essay (or more to the point a version that’s been polished up by a student who at least read the actual book) would probably get a C at UC Santa Barbara, perhaps a C- at UC Berkeley, and who knows, maybe an A at Columbia. But does this mean Copilot is now ready to steal our jobs? No, not yet. If it offered to ghost-write my next albertnet post, I’d politely decline.

Gemini’s essay

And how did Gemini, our final contestant, do? Here are some highlights from its first effort:

The irony primarily revolves around Clarice’s attempts to manage her mounting anxieties—namely her fear of a looming geography test and her worry about her best friend Betty Moody potentially moving away—while remaining oblivious to the humorous and poignant truths the reader recognizes.

The most potent source of dramatic irony is Clarice’s belief that she can control her own anxieties and the world around her by actively not confronting them. The geography test, which Clarice refers to ominously as “the Big Test,” is a persistent undercurrent of dread. She employs elaborate avoidance strategies, from organizing a “Worries Box” to distract herself with minor, manageable tasks. The irony lies in the reader’s understanding that avoidance only amplifies anxiety; the test is drawing closer with every page turned, yet Clarice invests her energy in delaying the inevitable.

The technique is equally effective in building the emotional tension surrounding Betty Moody’s potential move. Clarice views every sign of Betty’s packing or preoccupation through the singular lens of impending abandonment. The reader, however, is frequently given clues—through overheard adult conversations or context within the illustrations—that the situation is more mundane or less catastrophic than Clarice imagines. For instance, the boxes are not necessarily for a long-distance move, but perhaps a temporary one, or simply a major reorganization.

Again, the AI has written an essay that kind of sounds good and might fool somebody unfamiliar with the book, but it ultimately falls down because it has the basic facts wrong. Like ChatGPT, Gemini has no direct access to the full text and is just guessing (well, generating) based on the flotsam and jetsam that critics, reviewers, et al have strewn around the Internet. Gemini’s mistakes are not subtle. For one thing, there is no geography test mentioned anywhere in the book. Meanwhile, Clarice has a worries notebook, not a box, and she doesn’t spend undue time writing in it or employing avoidance strategies. Moreover, there is no scene in the book in which Betty packs boxes for her move, and the situation with Betty is not “more mundane” or “less catastrophic” than Clarice imagines: her best friend literally moves from London to San Francisco.

In the spirit of fairness, I gave Gemini another prompt to try to hone its response and focus on a specific scene of dramatic irony, and it came back with more of the same about the geography test and Betty’s move, adding, “The peak of this irony comes when the news of Betty’s situation is finally revealed, specifically that Betty is not moving away forever, but just temporarily to the apartment upstairs due to a family situation (often related to house renovations).” Completely false, and completely lame.

I pointed out all the errors and asked Gemini if that was really the best it could do. It replied, “The detailed essay you requested on Clarice Bean, Don’t Look Now was written using the Gemini 2.5 Flash model. If you were looking for even deeper, more complex analysis, the Gemini 2.5 Pro model (available with Gemini Advanced) would be the choice optimized for that specific level of complex, multi-step reasoning.” Oh, so if I want a better essay I have to pay for it? What cheek! I almost decided to just give Gemini an F and move on. That would have made this post shorter. But, doggone it, if I’m going to do this, I’d better do it right.

Gemini’s second essay

I switched to version 2.5 Pro (which turns is offered on a limited basis for free) and here’s the gist of its new essay:

The central irony is this: the very “spy” techniques Clarice uses to gain control and uncover the truth are precisely what blind her to it, generating both the novel’s humor and its profound sense of childhood helplessness.

This irony is sharpened in Clarice’s “investigation” of her parents. Overwhelmed by stress (which the reader understands is likely related to their friends’ move, finances, or just the chaos of family life), Clarice’s parents have tense, private conversations. Clarice, filtering these events through her Ruby Redfort lens, interprets them as “clues” to a secret, impending divorce. She misreads their mundane exhaustion as a sinister conspiracy. The dramatic irony is that Clarice’s hyper-vigilance, her constant search for meaning, makes her less perceptive, not more…

Ultimately, the book’s title, Don’t Look Now, serves as the most direct summary of its central irony. Clarice believes her “looking”—her spying and investigating—is the solution. But the reader knows she is refusing to look at the one thing that matters: the deep, painful, and very normal emotion of sadness. She avoids the emotional core of her problem by obsessing over its mechanics.

I have to say, this is much better. Version 2.5 Pro didn’t hallucinate, and its argument is more cohesive. It’s not perfect; for example, it misses Clarice’s epiphany about the limits of the Ruby Redfort book and thus overstates her lack of perception. But this essay so much better than what 2.5 Flash “wrote.”

So is that it, I give Gemini a C+ and move on? Not quite: remember, this chatbot benefited not only from my invoking its 2.5 Pro version, but from all the coaching I gave it in the chat. This distinction is crucial: AI LLMs do much better when you feed them high quality prompts and lots of feedback to supplement their training data. It’s equally important to understand that your input is not itself training data that the model can use going forward. The benefit you provide dies with your session. Thus, AI doesn’t learn and get smarter the way a human would; its progress is much more gradual. Which brings me to:

Gemini’s third essay

To see how 2.5 Pro would do without all the coaching, I opened a fresh session on my work laptop (i.e., totally different login, no history of my chats). (Don’t worry, I did this on the weekend.) (If you’re my boss reading this, congratulations on finding my blog, and please consider that my working knowledge of AI is surely valuable in the workplace and you should give me a raise.)

I guess I wasn’t surprised that 2.5 Pro didn’t do so well this time, but what did surprise me is just how badly it crashed and burned. Here’s an excerpt:

The plot is set in motion by a catalyst of deliberate misinterpretation. A cryptic, unsigned letter containing the vague warning, “something terrible is going to happen,” is received not as a piece of misdelivered junk mail but as a profound, personal omen… The humor is generated directly from this disparity; the audience … understands that the “terrible” event will be domestic, not devastating. The characters’ frantic preparations—installing locks, suspecting neighbors—are thus rendered as escalating absurdities, a performance for an audience that already knows the final act.

OMG, it’s the worst essay yet: total hallucination. There is no cryptic letter in this novel, no locks installed, no suspicion of the neighbors. I called this out, the chatbot apologized profusely for having accidentally based its essay on a different book entirely, and then it tried again:

The gap between perception and reality generates the novel’s central tension. While Clarice is hunting for evidence of international espionage, the audience is processing signs of a painful family separation. The “mysterious man” Karl meets is not a sinister agent, but, as the reader strongly suspects, his father.

Again, pure hallucination! There is simply no “mysterious man” in the entire book. I challenged the chatbot, asking how it gets its source material, both when a work is under copyright and when it’s in the public domain. Gemini explained that for public domain works its training data contains the full texts and also “the centuries of critical, scholarly, and secondary sources,” and for copyrighted works “is built from secondary sources … book reviews, detailed plot summaries, fan wikis, essays, and educational matters about the book.” So basically it’s amateur hour: the AI can’t really differentiate between, say, an esteemed college professor and a (gasp!) lowly blogger. As you can see this doesn’t always work so well. I’m going to give Gemini 2.5 Pro a D+.

As an aside that perhaps ought to be my thesis, I’d like to point out that the better AI gets at writing student papers, the worse off students—and the whole institution of higher education—will be. After all, the point isn’t for students to edify their instructors through their observations; the point is for the students to think and write for themselves. Yes, this is hard, but the right kind of hard, and through this struggle they ideally learn how to think and write, and can one day contribute in the realms of actual, non-student writing such as books, articles, or—worst case scenario—blogs.

Activity #2: original art

I’ve tinkered a lot with AI-generated art, usually to generate pictures to run at the top of my blog posts. It’s been pretty hit-or-miss; a picture which doesn’t stray into uncanny valley territory, or commit a major gaff like the wrong number of fingers on a hand, is all I’ve realistically hoped for. Today’s exercise is simple: I pitted the platforms against one another in the task of creating a picture for this post, featuring Clarice Bean. You can see the winner at the top, though you might cry foul: the art I ended up using is from Whisk, Google’s latest “experimental” imagine generator. I resorted to this new tool because I just wasn’t happy with the runners-up, as you shall see.

ChatGPT’s art

I asked ChatGPT, “Can you make a drawing for me of Clarice Bean reading albertnet on her tablet?” Not surprisingly, it mentioned the copyright and said, “I can’t generate or reproduce images of her or derivative works featuring her likeness” but offered to “generate an image of a cartoonish, freckled, red-haired girl reading a tablet, in the style of a children’s book illustration, but not resembling or referencing Clarice Bean specifically.” I agreed and here’s what it came up with:


I think you’ll agree that’s just about the most boring picture ever. It also has the classic issue of the subject holding the tablet backwards. This is just not that hard a prompt … what gives?

I said, “Make it a more realistic picture, please, and she should look a bit older, and have her in an armchair in her attic bedroom with a desk lamp, and reading the Ruby Redfort Survival Guide.” Maddeningly, the chatbot came up with a picture that was almost perfect, except that made her look a bit too old (about 15) and gave her Instagram-worthy boobs, which seemed inappropriate and unseemly. The picture didn’t show a lot of skin, but still … totally unusable (and I don’t even want to post it here because it’s in such poor taste). I replied, “Please make her a bit younger and flat-chested.” The chatbot chided me: “I can’t modify or generate an image based on physical or anatomical details like that.” Like it was basically calling me pervy! It even offered to “create a child-appropriate illustration,” as though I’d asked for something that wasn’t. Sheesh.

Copilot’s art

I gave Copilot the same initial prompt I’d given ChatGPT, and here’s what it came up with:


This is almost as boring as ChatGPT’s picture, and for some reason it looks faded and I couldnt get Gemini to fix that. At least the tablet is facing the right way. Note that Clarice is wearing the same red-and-white-striped shirt in this picture as the ChatGPT version of her, which is curious given that such a shirt appears nowhere in any of her books (at least that I can find). It’s actually the shirt Waldo wears, which I’d prove to you if I could only find him.

Other similarities of this art include the hair being the same length, the art having the same level of detail (barely more than a cartoon), and a complete absence of any details in the background. In delivering the picture, Copilot said, “Here you go - a stylized, collage-like illustration of a child reading a tablet, inspired by the playful textures you mentioned.” I don’t know what it means by collage-like, and I didn’t mention any “playful textures.” Whatever, chatbot.

Gemini’s art

I gave exactly the same art to Gemini, and it produced the corniest, least aesthetically pleasing picture yet:


Obviously this is a matter of taste, but would you agree there is no charm here? And what’s with the red-and-white-striped shirt appearing here, too? What are these AIs keying off of?

In Gemini’s defense, at least the little thought bubbles bear a slight resemblance to some of the art in the actual book. But again the tablet is backward and “albertnet” is spelled “alphabertnet” (weird misspellings being a common screw-up with AI art).

Frustrated by not having any good art yet, I tried ImageFx, another Google AI tool, and it gave me a photo-style picture with lavish detail, featuring both Ruby and her brother rocking red-and-white-striped shirts. I think it’s some kind of global AI conspiracy. What a relief when Whisk broke the cycle and generated the worthy picture you saw at the top of this post. I particularly like how Clarice is kind of staring off into space instead of at the book, clearly either pondering what she’s just read or distracted from her book by all the difficulties she’s working through.

Well, at long last that’s it for today. Tune in next week because I plan to pitch these chatbots against one another again, this time writing poems in dactylic trimeter based on the best prompt an AI was ever given.

Other albertnet posts on A.I.

—~—~—~—~—~—~—~—~—
Email me here. For a complete index of albertnet posts, click here.

Saturday, March 29, 2025

What Is ChatGPT Great At (and Not)?

Introduction

If you are reading this post long after its March 2025 publication date, you might become puzzled at its many failings until you realize, “Oh, wait, this was written back when Dana wrote his own blog posts instead of assigning them to GPT-28-turbo-XL-prime! The lameness is because he did his own very light research instead of basing his observations on the entire body of knowledge of the Internet, and because it’s a plebian human voice instead of an infinitely exalted and witty A.I.!”

I have now blogged 14 times about A.I. and its evolution. My last ChatGPT check-in was about three months ago. The A.I. version hasn’t changed since then; as of this writing it’s still GPT-4-turbo. But what has changed is the range of tasks I’ve experimented with. I now realize my previous posts failed to appreciate some of the things ChatGPT does really well. This post showcases those, while also providing commentary on what the A.I. still does not excel at (and likely never will). You’ll also learn more about why you may have seen an annoying banner about cookies at the top of this blog.


Caveat

This post is mostly about ChatGPT though it touches on Google Gemini. What it doesn’t cover is the “Visual Look Up” feature on Apple’s iOS platform that leverages “Siri Knowledge.” I don’t currently own any Apple products (except an iPod mini in a drawer somewhere) so all I know about Siri is that it did a comically poor job of identifying the breed of my brother’s cat today, based on this snapshot he sent me:


How an A.I. could think any image looks like both a cougar and a wallaby is beyond me. I’m going to assume Apple is so far behind in the arms race that we can simply ignore it for now.

Real-world problem solving with GPT-4-turbo

Until recently I’d only messed around with GPT-4 for the purpose of evaluating it (and, whenever possible, mocking it). But then I hit upon a real-world use case and dove back in. My motivation, which I’m sure you’ll relate to, was: HOT CASH MONEY. Who wouldn’t want this, other than those tedious killjoys who spout aphorisms like “Money is the root of all evil”?

By way of background, I’d noticed that the albertnet page view count had soared in recent months. It took this blog something like 14 years to reach a million page views, but in the last six months alone I’ve now seen almost 1.4 million more. But then, isn’t this how the Internet works? Moore’s Law? Nielsen’s Law? All that compounding magic? In the whole time I’ve had this blog I never even considered monetizing it through ads, but every man has his price. (I’m not sure exactly what mine is, but I reckon I’ll know it when I see it.)

Driven mad with money-lust like one of the guys in “Treasure of the Sierra Madre,” I needed answers—fast. So I asked GPT-4-turbo, “My blog, www.albertnet.us, has received 1.2 page views in the last three months and traffic is increasing. If I turned on Adsense, approximately how much money would I earn per month?”

Yes, “1.2 page views” is a typo, but I didn’t make it here … that’s actually what I asked ChatGPT. It replied, “What the hell do you mean 1.2 page views? How do you have 2/10 of a page view? Did some user barely see the screen, like out of his peripheral vision? Or are you just whacked out on coke and smack and typed your query wrong?”

Okay, you got me … that’s not at all how GPT replied, though honestly I think that would be the better answer. What it actually provided was a lengthy essay, full of data points and computations, answering this useless question. My favorite part of the response was, “Number of Page Views (Traffic) – You’ve mentioned you have 1.2 page views in the last three months, which is approximately 400 page views per month (assuming the traffic is consistent).”

Huh? How do you get 400 by dividing 1.2 by 3? I guess the chatbot arbitrarily assumed the figure I provided was in thousands. That’s a pretty big logical leap, and GPT didn’t document the fact of this assumption. It then proceeded to run a bunch of calculations based on 1,200 views, the punch line being that I could make about $2/month. So the more succinct answer would have been, “Dream on, bloggy-boy.”

When I corrected my original query to 1.2 million page views, GPT-4-turbo reran its calculations and informed me that I might expect to earn something in the neighborhood of $2,000/month in passive income. Now we’re talking! It did suggest a number of caveats, such as how my  results might be affected by the geographical location of my readers, the positioning and type of ads, ad targeting, how well ads match my content, user engagement, and so on. I asked it a bunch more questions specific to Adsense, whether GPT’s estimated click-thru rate (CTR) assumption is realistic, etc. While it provided all kinds of useful info, it missed one very important rule of thumb: if something seems too good to be true, it probably is.

I mean, come on … albertnet is a blog about nothing. I’m not going on political rants that cause trolls to leave endless acerbic comments and then forward my post to 90 friends with an exasperated preface like “can you believe this shit?!?!!?!” If all it took to create a nice passive income stream was to blog every single week for 15 years straight so that after more than 3,500 hours of writing you’ve amassed over 750 posts, comprising over 2 million (juicy, searchable) words, then everybody would be getting into this business, obviously. If quality, rather than nudity, attracted people’s attention, every liberal arts grad on the planet would be driving a Benz. (Well, except for me, because regardless of my income—actual, theoretical, or pipe-dreamed—I will always be the world’s cheapest man driving a used Volvo.)

I feel really bad for the earnest blogger who sees all this traffic growth, does a basic ChatGPT query, thinks he can trust the response, and makes a lot of effort adding ads to his blog to harness this new fountain of riches. I hope nobody is that naĂŻve. Since I’m not, my first impulse was to get a second opinion. So I put my query to Google Gemini, without the typo this time, and it gave me a very similar answer: I could make right around $2K a month, just for setting up Adsense and then sitting on my ass!

This seems like the kind of claim I’d get from a spammer. In fact, I’m going to check my spam right now to look for this kind of hoax. Hmm, that’s odd. Though 22 of my 50 spam messages are from recruiters, none is this specific. I do, however, have a message from “Super-size” titled, “Now imagine each night, having 5 or 10 concubines around you, each one craving your masculine essence in them. #632352.” This subject line is interspersed with various emoji including, oddly enough, an avocado. Is avocado a concubine’s favorite food? Let me ask ChatGPT. Okay, it replied in the negative, pointing out that concubines were prominent “in ancient or medieval times, when avocados were not available in their regions.” I think the innuendo of “in their regions” was accidental. (And now I’ve realized how long and pointless a digression this has been. I’m tempted to apologize, except this might end up being my favorite paragraph of this entire post.)

The caveat GPT-4-turbo should have provided is, “My calculated revenue assumes the page views are from actual readers, not bots.” The idea of bots grossly polluting my page view stats was my natural assumption, but not one GPT addressed. I think this is an important object lesson: it doesn’t matter how useful A.I.’s responses are if you don’t know to ask the right question. Perhaps A.I. will advance to the point that it would not only sanity-check my page view stats, but would be the one to keep an eye on my blog traffic to watch out for moneymaking opportunities in my stead. (If and only if it knew albertnet to be an amazing viral sensation for reals.)

Where ChatGPT shines

After feeding me all that false hope, ChatGPT asked if I’d like help setting up Adsense on my blog. I decided instead to enlist its support vetting the quality of the page view stats. Having drilled down a bit on my own (which Blogger doesn’t make super easy, by the way), I discovered that page views from France were 12% of my total over the last six months, 14% over the last three months, and 39% over the last 30 days. I shared this with the chatbot and queried, “Is there A.I.-driven or bot type traffic that would originate in France that could artificially inflate the metrics around my readership?” (I now regret the specificity of this, as I was clearly “leading the witness.”)

ChatGPT responded with a clear and detailed essay about the probable causes, including “Bot Traffic (Most Likely Explanation).” It filled me in about scrapers and SEO crawlers and A.I. training bots, and suggested I use Google Analytics to investigate further. This ended up being an excellent suggestion and led to my most engaged use of GPT ever.

First I quizzed the chatbot about whether Google Analytics has a free version (it does), whether I’m giving up any privacy (basically not), etc. Then I set up Google Analytics (aka GA4), which was pretty straightforward, except I noticed in the Terms of Service that I’m expected to comply with GDPR (the EU General Data Protection Regulation) when gathering this detailed user data. I happen to know what GDPR is, so I asked GPT all about it, in terms of what I’m really expected to do to comply. It turns out that compliance is a royal pain in the arse (my words, not ChatGPT’s). Since I do get readers from Europe (whether it’s 39% of all traffic or not), I need to have a publicly posted privacy policy and a banner announcing my use of cookies (which is how GA4 can track usage). I almost abandoned the whole project, on the mere assumption that my page view stats are so obviously bogus I don’t need to expend all this effort verifying it, but then … what if these traffic stats aren’t bogus? What if I really could just sit back and rake in money? Isn’t it worth spending some time and effort investigating the possibility?

I asked GPT for some nice boilerplate text for the privacy policy, and though much if its response was unusable, some of it was good, and if nothing else this rough draft prevented writer’s block and paved the way for my policy, which you can read here and which I’ve linked to in my blog’s footer. (I’ll need to revise that policy pretty radically, as you shall see, but it’s probably a good thing to have anyway.) The harder task was creating that cookie banner, since it’s not just a static digital placard but an actual functional utility that captures a user’s cookie preferences and turns them into policies that impact the behavior of GA4. That is no small feat, and probably nothing I’d tackle on my own.

Before I pressed on I had a long, rambling discussion with ChatGPT about how to get everything going. I learned a ton, including info about the following:

  • The various metrics I’d be able to get from GA4 (i.e., is this truly worth it?)
  • What free utilities exist that could be leveraged for setting up the cookie banner and how to choose the best one
  • Approximately how long it would take to set up the banner based on the chosen utility
  • How to create a Google Tag and write an HTML script for my blog’s template that would invoke it
  • How to create the HTML script that would invoke the banner
  • How to pause GA4 if I have trouble invoking the banner (spoiler: I did)
  • How to back up my blog’s HTML template before messing with it (though GPT didn’t suggest this, which again illustrates the difference between a) being able to describe how to do something, and b) taking the initiative to do that thing)
  • How to debug my script and figure out why it’s not working

These weren’t just general instructions it provided that I’d have to suss out on my own. GPT4-turbo provided sample script text that actually worked (eventually). Here’s an example of its suggested script:

 <head>
    <!-- Your other head tags -->
    <script async src="https://www.googletagmanager.com/gtag/js?id=YOUR_TRACKING_ID"></script>
    <script>
        window.dataLayer = window.dataLayer || [];
        function gtag(){dataLayer.push(arguments);}
        gtag('js', new Date());
        gtag('config', 'YOUR_TRACKING_ID');
    </script>
</head>

To reiterate, I am not a seasoned HTML jockey and would have struggled with this syntax, to say the least, were it not for the chatbot’s help. And even if I had originally built my blog from scratch (i.e., coding all the HTML myself without a Blogger template), I’d have been rusty enough now that I’d have been wise to leverage GPT for this task anyway. As I went through all this scripting, it dawned on me why a lot of people are worried about A.I. taking our jobs. This is just basic HTML but GPT was hugely helpful; If I were a full-on programmer and suddenly became (say) twice as efficient because I was grabbing blobs of basic code for simple operations instead of creating them from scratch every time, I’d naturally consider how all my colleagues have also become twice as efficient, and I’d start to worry about my employer realizing they could make do with half their programming staff. Scary stuff.


The upshot

Once GA4 was up and running, ChatGPT was very helpful in walking me through understanding all its metrics, not all of which were very intuitive. In a perfect world, I’d have discovered an average engagement time of ten minutes per post, indicating actual human readers. In reality, I learned that—guess what?—average engagement time is under three minutes, and GA4 shows way fewer page views than the Blogger stats. In other words, Blogger most likely is reporting on a lot of bogus visits from bots. If ChatGPT were like a really cool know-it-all big brother, it would have said, in response to my very first inquiry about the growth in traffic, “Dude, don’t trust the Blogger stats. They’re useless.” I wouldn’t have had to do all this research.


After asking ChatGPT a bunch of questions related to the delta in traffic as reported by each platform, I had it recalculate the ad revenue I might hope to get from my blog in light of the better data. It estimated about $0.40/month, and then went on to suggest a whole bunch of ways I could improve user engagement. I then led it on a thought exercise about how much of the real traffic is based on old posts, since a) albertnet posts are not timely, and b) the longer a post is up, the more views it will gradually accrue. GPT agreed with my assessment: that any improvements going forward would only marginally increase traffic and engagement, as they’d only apply to new posts.

Next I asked GPT for its best guess as to how much improvement I could achieve if I implemented all its suggestions … double? triple? tenfold? It replied that “doubling or tripling engagement is probably a reasonable and achievable short-term goal.” (This doesn’t impress me as intelligent … I think the chatbot is just highly suggestible.) I went on to ask, “Do you think it’s worth implementing these strategies with the goal of monetizing my blog through ad revenue?” It provided another long essay that concluded, “Yes, but with realistic expectations … treat ad revenue as a potential bonus or passive income stream, and consider other monetization strategies as well (affiliate marketing, sponsored content, or selling your own products/services).”

And this is where, I think, A.I. is still falling short. It’s great at helping the user with nuts-and-bolts technical tasks, especially those of the type performed countless times by other users (for example, inserting scripts to invoke GA4 and/or a cookie banner). But synthesizing a lot of information and drawing the best conclusion is still beyond its ability. By its own reckoning, my real, human traffic would bring in $0.40/month, and by implementing all its suggestions I might triple user engagement … but it failed to grasp that earning a mere $1.20/month isn’t worth any amount of effort. A.I. was ultimately unable to suggest the right strategy for me to take regarding my blog.

One last thing…

The sad part of this tale of exploration is how useless all my effort has ended up being. If I had any reason to suspect that, after fifteen years, my blog would suddenly go viral, I could keep an eye on Google Analytics to savor my success … but I don’t. I might as well be a frog looking in the mirror every morning to see if I’ve miraculously become a prince. Not that I actually care, mind you … as described here, I’m happy to be a humble frog croaking out my unsung song. But there’s no point bothering new readers with that cookie banner, especially since—as I recently discovered—the damn thing doesn’t even work.

This is another thing ChatGPT overlooked … it failed to suggest that I actually put that banner through its paces, which I’ve now done. Through basic experimentation I’ve discovered that it doesn’t end up mattering what preferences the user selects … his session is duly recorded in GA4. (Only if a user uses a Private or Incognito window are his sessions ignored … even if he allows all cookies.) Meanwhile, site visits from mobile users are not counted at all by GA4, I have just determined. That may be because I never got the banner to work on mobile, and Google can tell this and wants to observe the GDPR rules.

So now I have to go shut the whole thing down, to maintain GDPR compliance. The entire exercise was (to borrow from Shakespeare) “the expense of spirit in a waste of shame.” (Shakespeare was writing about lust, but I think his sonnet also covers the lust for money quite nicely.)

Check back in a week or so and (with ChatGPT’s help) I’ll have backed out the GA4 scripting and gone back to an unstudied, non-monetized blog with no banner. My privacy and cookie policy will have had a makeover as well. I’m no richer for this little exercise, but a bit wiser, and now you are too.

Other albertnet posts on A.I.

—~—~—~—~—~—~—~—~—
Email me here. For a complete index of albertnet posts, click here.