Showing posts with label chatbot. Show all posts
Showing posts with label chatbot. Show all posts

Saturday, November 15, 2025

More AI Smackdown - ChatGPT, Copilot, & Gemini Write Poetry

Introduction

Two posts ago, I described what I think is a fundamental dichotomy between two central capabilities of modern AI chatbots: 1) helping with a nuts-and-bolts operation like coding software or scripting HTML, and 2) creating something original, like an essay or story. The first category involves being a resourceful researcher blessed with excellent natural language processing; the second is probably closer to what humans are (so far) uniquely capable of doing.

Earlier this year I did a whole post on the first category, “What is ChatGPT Great At (and Not)?” And last week I blogged about one aspect of the second category: writing a scholastic essay. To further explore AI’s ability to generate meaningful content, and to evaluate its ability to truly understand language, I turn this week to poetry. That is, I decided to have the three dominant chatbots—Gemini, ChatGPT, and Copilot—write a poem in an unusual meter: dactylic trimeter, a poetic form I learned in high school (details here). I chose this meter because, as described here, ChatGPT does a pretty good job at the classic Shakespearean sonnet in iambic pentameter, but I wonder if that’s just really good parroting since there’s such a vast amount of training data out there for that. I think this exercise really puts the chatbots through their paces, giving us insight into which is the closest to being truly intelligent. As you shall see, the differences in performance are not subtle.


(Custom art by Whisk. No rights reserved.)

Gemini’s effort

To start out, I quizzed Gemini about dactylic trimeter, to see if it knows what I’m even talking about. Gemini correctly stated that the rhythm of such a poem would be “DA-da da | DA-da-da | DA-da-da,” and an example it created of the form was reasonably close. So far so good. But then, to make the rhythm better, I instructed  the chatbot to add an extra trochee at the end of each line. A trochee is a two-syllable word with the stress on the first syllable, as in the word “praises” and the word “spirit.” As an example of this modification to dactylic trimeter, I provided Gemini these lines (that I took from a poem you can read here, in this albertnet post):

Once in a while a voice will sing praises,
Something to levitate everyone’s spirits.

A really smart AI, I would argue, could reverse-engineer the meter from those two lines alone, but I went one better and described exactly what I wanted in technical terms. Gemini correctly stated that the rhythm would therefore be “DA-da-da | DA-da-da | DA-da-da | DA-da” but its initial attempts at it were totally screwed up. I gave it a lot of coaching. I guess this is okay; a human with actual intelligence might require this as well.

Moving on, I prompted Gemini, “Now I would like to see if you can write such a poem based on an essay I provide. You can work in as much as you think works, understanding that not everything will fit. If I paste in a 1,600-word essay could you do that?” It said (in its usual obsequious way), “That sounds like a very interesting challenge! Yes, I can absolutely take a 1,600-word essay and distill its key themes, concepts, and emotional tones to create a poem in the specific Dactylic Trimeter + Trochee (11 syllables) structure.” I added that the poem should be ten lines long, with an ABBA rhyme scheme for the first two quatrains and then a rhyming couplet at the end. And I told it, “Focus on the idea of caprice: choosing the more difficult route simply to exercise one’s radical freedom and prove one’s mettle to oneself.”

Now, if you are a devoted follower of albertnet, you may recall that some years ago I wrote a poem in dactylic trimeter about my favorite cycling hill climb, the one-two punch of the steep, twisty roads El Toyonal and Lomas Cantadas. And, as with most of the  poems on this blog, I provided copious footnotes and commentary, just as you’d get in an anthology of Shakespeare’s sonnets. To help AI write a good poem, the “essay” I uploaded was the compilation of my footnotes, so that Gemini had plenty of material to work with: the fleshed-out observations, feelings, etc. behind my Lomas poem. The idea here is to give you, reader, the ability to do a side-by-side comparison of two poems on the same topic—one written by your amateur blogger, the other by a much-ballyhooed AI. You can decide for yourself who did the better job. Here is my poem (and you can click here for the footnotes):

Ode to Lomas Cantadas

El Toyonal is a beast of an uphill;
Pedaling here is a bit self-defeating.
Lomas Cantadas will sweeten your beating:
Insult to injury, sweet overkill!

Most riders wisely take Wildcat Canyon.
Half the uphill, after all—and you’re tired!
Wisdom, alas, is a flaw when you’re mired
In glory, in notions of being a man.

Sometimes caprice is much better than ease:
Radical freedom is here to be seized!

Here is Gemini’s poem, representing its effort to incorporate all my footnotes and commentary:

Reckless Choice

The simple road beckons And calmly the wind calls,
We stand at the turn now, Deciding to stay free.
To choose the hard pulses Our own perversity.
We turn for the suffering feeling the high walls.

The logic demands the Less burdensome track-way,
But we must display the Radical self-might.
We take the bad option Defying the pure light.
To prove we are strong to Our own confidences.

Yet this is the price of The joy and the pain-fall,
A reckless decision That answers the soul’s call.

So … how would you rate Gemini’s poem? Is AI going to come after our poets and put them out of work? Frankly, I’m less than impressed. The meter is a crude approximation of what I asked for, but as you surely felt (whether or not you could articulate the precise problem), almost every line is too long by one syllable, and instead of the easy, rollicking rhythm of a proper Krooprian poem, Gemini’s ode staggers and stutters. The AI pretty much blew it.

In case you really want to geek out, here’s exactly where the AI went wrong: in nine out of the ten lines, for no apparent reason, it preceded the first dactyl with an unstressed syllable, violating the rule I set forth about dactylic trimeter. Meanwhile, words like “perversity,” “radical,” and “confidences” aren’t padded out properly so we either have to mispronounce them to preserve the meter, or sound them out naturally which disturbs the rhythm—and since we don’t know which to do, we’re stopped in our tracks.

As far as the content, you can sense the AI compromising sense in order to (try to) achieve the rhythm and rhyme. Why else would it use the word “pulses,” the phrase “high walls,” and the Franken-words “track-way,” “self-might,” and “pain-fall”? The chatbot is practically babbling.

So is that it? I give it, say, a D and we move on to the next AI? Not so fast. I made the rookie move of feeding my prompt to the default version, 2.5 Flash, instead of the flagship version, 2.5 Pro, which Gemini tells me is “more adept at tasks requiring complex reasoning, deep instruction following, and high-quality creative output.” The good news is, as the AI assured me, all the context of my chat with 2.5 Flash can be used by 2.5 Pro, so it was super easy to switch to Pro and have the chatbot try again. Here’s its second effort, using the 2.5 Pro version:

The Gauntlet Thrown

Sober thoughts point the logical straight-way,
Turning left proves a man is still willing,
Recklessly choosing what he must still be,
Turning from all that the reasonable say.

This is the freedom to throw all the rules about,
Knowing the payment and what it will cost him,
Twenty hard minutes of everything lost now,
Heeding the irrational gonzo-man’s shouting.

A purely perverse act of personal pride,
Proving the mettle he carries inside him.
Well, I suppose it’s a bit better. Only one line erroneously starts with an accented syllable, and three of the lines actually have the correct meter (a modest improvement over 2.5 Flash’s zero). There’s some nice alliteration in the penultimate line. I also appreciate that 2.5 Pro didn’t do weird stuff with capitalization in the middle of each line. But the rhyming is pretty janky, and we still have these weird bolted-together phrases “straight-way” and the truly bizarre “gonzo-man.” To top it off, the very last word of the poem violates the rhyming convention, so the poem goes out on a flat note. This second poem is better but not by much.

Conclusion: if Gemini offers to write my next albertnet ode for me, I’ma be like, “Naw, I’m good.”

On to our next contestant.

ChatGPT’s effort

I asked ChatGPT if it could write a poem in dactylic trimeter. It said, kind of curtly, “Definitely — I can write in dactylic trimeter.” Unlike Gemini, it didn’t bother to give me a sample straightaway. Maybe this is to save on computing cycles (and thus energy and water), which I can appreciate. But ironically, this is the one chatbot whose capability I have reason to doubt, because as you may recall, I tried this prompt before, almost verbatim, with version GPT-3.5, back in February of 2023. That chatbot crashed and burned with this poetic abomination:

Pedaling up Toyonal, heart pumping fast
My mind is clear, free from the world’s contrast
With caprice and freedom, I take each turn
Pushing through pain, for what I will earn
Onward I climb, up Lomas Cantadas
My legs scream out, but my spirit never fades
The wind in my hair, the world fades away
A moment of peace, that I long to stay
And at last I reach the summit high
A sense of triumph, that no words can deny

Ugh. I almost threw up in my mouth rereading that. As you can see, GPT-3.5 ignored my instructions about meter, made a very half-assed effort to weave the theme of caprice into its poem, and worst of all produced pure treacle, like the verbal equivalent of Thomas Kinkade. Needless to say I was eager, this evening, to see how the more modern version would do.

I gave GPT-5.1 the same prompts as Gemini, uploading the same 1,600-word essay, and here is what it produced:

Lomas Caprice

Turning toward Toyonal, reckless and laughing
Choosing the steep when the wise path is brighter
Mocking my reasons, the prudent, the lighter
Savoring folly my sane side is quashing

Freedom is choosing the climb that will break me
Just for the thrill of defying the sensible
Proving I’m bent for the wholly nonsensical
Riding the harder way just ’cause it shakes me

Pain is a dare that I answer with pleasure
Freely embracing the hurt as a measure
Well, ChatGPT has come a long way for sure, but GPT-5.1’s effort is only somewhat better than Gemini’s. Certainly the meter is better, with a majority of the lines being correct. But the content is really off, with a bunch of the words clearly chosen just to satisfy the technical requirements without adding much meaning. The bit about “wise path is brighter” really makes no sense and is clearly just there for the rhythm and rhyme, no more sophisticated than Hall & Oates’ “your kiss is on my list.” In the next line, who is doing the mocking? And how does “the lighter” fit into anything? Lighter sky? Lighter weight? Cigarette lighter? It’s just a random word dropped into the poem. And in the next line, the word “quashing” in no way rhymes with “laughing” and doesn’t make sense as an intransitive verb. (“What are you doing this weekend?” / “Oh, you know, I’ll just be at home, quashing.”)

 I confess, I rather like the line “Freedom is choosing the climb that will break me,” but then the poem loses momentum again and commits rhythm-sucking metrical errors on the next two lines (though I like “bent”). The eighth line, suggesting that a hard climb “shakes me,” is lame, another word selected only because it rhymes. And that last line? “Freely embracing the hurt as a measure”? Huh? What is it measuring? This poem is lame.

Since AI does its best work when you iterate with increasingly refined and specific prompts, calling out what it did wrong in its previous attempt, I decided to give ChatGPT another chance, and told it, “I think it would be better if it didn’t assume what you and I know already about this climb. Consider that somebody encountering this poem for the first time wouldn't know that Wildcat Canyon is the easier climb, and that choosing the 1-2 punch of El Toyonal and Lomas Cantadas makes no logical sense but appeals to one’s love of suffering and sense of caprice. So, please try again on the poem and give the reader enough background to grasp all this and thus to understand the choice.” It came back with a poem that was quite broken, with the same issue that Gemini’s first effort had: starting each line with an unstressed syllable. It also screwed up the rhyme in the second quatrain. I coached it repeatedly to fix these issues, and after several tries this ended up being its best effort:

Reckless Climb

Climbing the hills of green Berkeley foothills,
Pedaling hard as the thighs start to quiver,
Wheels weaving wild like a paperboy’s river,
Lungs heaving fire as the body fulfills.

Turning to torment, no reason persuades me,
Pain blooms in muscles yet joy is commanding,
Twists of the road, and the thrill never fades me,
Searing the legs, but the spirit is standing.

Pleasure is folly, the wholly absurd,
We choose what will hurt us, yet laugh at the hurt.

Right off the bat, the first line has three problems: it trips us up with a missing syllable; the hills are not always green; and hills/foothills is somehow both redundant and oxymoronic. The rest of the lines have correct meter, but a bunch of really clunky constructions. Wheels weaving wild? Really? And what could “like a paperboy’s river” possibly mean? “Lungs heaving fire”—what am I, a dragon? And the body fulfills what? (I just asked ChatGPT if “fulfill” can be an intransitive verb, and it replied, “No — “fulfill” is fundamentally a transitive verb, which means it requires a direct object.”) The mention of the twists of the road isn’t connected to anything, and “thrill never fades me” makes no sense—I mean, is the thrill like the sun, and I’m a curtain? There’s just an abundance of nonsense here, though I do like the last line. My last complaint is that the AI didn’t consistently follow the rhyme scheme I asked for: its second quatrain is ABAB, not ABBA. I asked it five times to fix this and it never did.

In summary, I’d say this poem is marginally better than Gemini’s effort (I mean, it doesn’t have anything as glaring as “irrational gonzo-man’s shouting”) but it’s basically unusable. Maybe GPT-6 will do better.

Copilot’s effort

Copilot knew exactly what dactylic trimeter was, and its initial example was remarkably good. I suppose it could have plagiarized something, or has a bit of boilerplate at the ready, but in any case I was impressed.  I gave it the same extra instruction about tacking on an extra dactyl, and it got it almost right on the first try; a bit more coaching and it nailed it. After invoking Smart (GPT-5) mode, I fed Copilot the full prompt, with the same footnotes from my Lomas Contadas poem. Here is what the chatbot came up with:

Caprice on the Hills

Turning away from the sensible roadway,
Madness compels me to suffer the steepness.
Reason protests, but I laugh at its weakness,
Choosing the torment that beckons me boldly.

Freedom is found in the reckless decision,
Mocking the logic that urges avoidance.
Pain is a proof of my will and its poignance,
Glory is born of capricious derision.

Taunted by Toyonal, daring and cruel,
Lomas Cantadas declares I’m no fool.

Wow. Just wow. It is so much better than what Gemini and ChatGPT came up with. The meter is actually perfect. (Full disclosure: the meter falls short of what’s in my poem, but only because I forgot to give the AI one more instruction—to shorten the final line of each quatrain, and both lines in the couplet, by one syllable, to give them more punch. To be honest I’d forgotten I did this with my ode.)

I think the AI did a pretty job with the content as well. Calling my will “poignant” is a bit of a stretch perhaps, but we could chalk it up to poetic license. Meanwhile, the word “roadway” (like “poignance”) is nonstandard and evidently chosen to fit the meter, and also doesn’t rhyme very well with “boldly,” but this is nitpicking. My main issue with this poem, which is the same issue I had with the others, is that it doesn’t give the reader enough context to really understand some of the basics: that there’s an easy option (Wildcat Canyon) and a harder one (El Toyonal + Lomas Cantadas).

Perhaps as a victory lap, Copilot offered to write two more versions with slightly differing tones (the last being a drinking song). These weren’t actually as good as the first so I won’t bother dissecting them, but I did like the final couplet in one of them. I decided that since Copilot’s initial poem was good enough to warrant refining, I would point out its central error and let it try again (using the new couplet). Here’s its revised version:

Caprice on the Hills

Wildcat is gentle, a steady road climbing,
Toyonal rises with crueler intention.
Still I turn left in a rash intervention,
Mocking the voice that would counsel good timing.

Freedom is folly, a dare to my reason,
Choosing the torment that makes no good sense.
Pain is a jest and its cost the expense,
Glory is born of capricious decision.

Madness is mettle, a jest I renew,
Lomas Cantadas — I suffer for you.
Again, it’s not perfect. The bit about “turn left” doesn’t really make sense without more context. And what is meant by “counsel good timing”? There is no timing here. In addition, the phrase “its cost the expense” seems redundant. But those matters aside, the meter is solid, the rhyming consistent, and that final couplet is a real banger.

These AI chatbots always seem to want to extend the dialogue and provide more and more and more, which is kind of a double-edged sword. On the one hand, as human beings we should always be working to limit our time online and get out there in the world, right? On the other hand, refining what we get from chatbots is pretty key to making them an effective tool. So when Copilot asked if I’d like it to craft a prose introduction to the poem, I suddenly had another idea: what if I asked it to now create its own footnotes? This post is long enough already so I won’t post them here, but let me say that Copilot did a pretty good job on that.

And here is where I see this AI having a role with a real human writer (at least at the student or blogger level): it could probably help with writer’s block simply by producing something worth polishing. It kills me to concede this, actually, and I am far too proud to ever resort to this kind of “Hamburger Helper” approach to my own writing. But honestly, a cyclist who would like to compose a ride-themed poem in dactylic trimeter, replete with footnotes, could do worse than to start with Copilot. (Neither poem above truly passes muster, but taking the best of each, and from perhaps a few more attempts, and then replacing all the weak parts with our own lines, would be easier than—albeit still inferior to—starting from scratch.) The output of such an exercise might actually have some value, versus the writer getting frustrated, giving up, and producing nothing.

Crucially, the thing the AI will never be able to do is go on the bike ride, have that experience, and grasp what is important about it. So a human could start there and then get some help from AI in expressing himself or herself, since not everyone has the luxury of a liberal education. If AI is called upon to bridge that gap, the current Copilot is far better poised than Gemini and ChatGPT, I think we can now conclude.

If you read my last post, you may recall that Copilot did the best job of these three chatbots at writing a scholastic essay as well. Keep an eye on this one … Microsoft, through its partnership with ChatGPT’s OpenAI as well as its own resources, seems to be ascendant.

Other albertnet posts on A.I.

—~—~—~—~—~—~—~—~—
Email me here. For a complete index of albertnet posts, click here.     

Friday, October 31, 2025

Tech Reflection - Two Sides of AI

“This Halloween, I’m dressing up as generative AI. I’m going to show up to the party without a costume and just start stealing pieces of other people’s outfits.”
An X dispatch my niece screenshotted for me

Introduction

Is AI the amazing new technology that’s changing the world, or a petty thief that just steals people’s ideas and passes them off as its own? Does it actually carry out anything approaching thought, or is it just a zombie, stalking humans’ digital relics and muttering “brains … brains … brains” as it angles to get a piece of us?


In this post, I examine the two most fundamental functions AI chatbots can carry out, and draw a distinction between the two. I believe this can give us useful guidance in deciding how we ought to use this game-changing technology.

Ecclesiastes vs. Barthelme

AI is evolving fast, perhaps faster than our ability to understand it. I’m having to adapt; for example, I’ve stopped spelling it “A.I.” because leading media outfits like The New Yorker and The New York Times have now eschewed the periods. So if you’re reading this on your phone in a sans serif font you may have initially thought I was writing about there being two sides of Albert or Alfred. I asked ChatGPT what to do about this ambiguity between a capital “i” and a lowercase “L,” and it suggested I could “kern or tweak the glyphs.” I’m not exactly an expert at kerning glyphs, so I asked the chatbot how. It gave me all kinds of strategies, the best one for my blog format (HTML) being this:

AI <!-- default -->

A<span style="letter-spacing:0.05em;">I</span> <!-- slightly looser -->

So you can see, GPT is right there with an answer when queried about a technical operation that has been done before. But what about doing something creative and original? This is a fundamental distinction and I am going to propose we look at AI from two largely separate perspectives, for which I’ve invented labels:

  • Operational mode – I thought about calling this Ecclesiastes mode, for “no new thing under the sun.” This mode is about helping with a nuts-and-bolts operation (e.g., HTML scripting, DNS routing) that somebody else, probably many people in fact, already figured out and documented out there on the Internet for AI to gobble up, distill, pretty up, and present. Here, AI is basically a really good large language model that excels at combing through gobs of chaff to find answers, and organizes and summarizes information very clearly. I wouldn’t say it’s as parasitic as what’s suggested by the X epigram above, because lots of people freely post technical stuff to the Internet just to be helpful, without thinking of it as sacrosanct intellectual property.
  • Creation mode – I think of this as Barthelme mode, named for the writer Donald Barthelme, because I think he’s the epitome of totally original, wacky, one-of-a-kind creative types with an absolutely distinctive voice. In other words, this is the intelligence that I am quite convinced AI could never approach. By creation I mean using the full capability of your own mind to advance ideas that are new, and yours.

The trouble is, many people don’t make any distinction between these two general areas of AI, so on the basis of its prowess as a natural language search engine, they are be led to believe it can do a perfectly good job at creation mode. And since most people aren’t English majors, and in fact don’t respect English majors, AI platforms get to roll out some pretty inferior writing and everybody thinks it’s genius. (This widespread lack of sophistication is also why McDonald’s makes so much money.)

So what?

For many years, as I’ve lamented at length in these pages, kids have been told all the jobs are in tech, and they need to study STEM. And now, many of the kids who dutifully followed these marching orders are graduating from college with Computer Science degrees and not getting jobs, and tech is laying off gobs of people. Next time I meet a STEM major I’m gonna ask him, “Computer Science? What are you gonna do with that?”

So how did STEM go from meal ticket to a food stamp? Well, I think it’s largely because AI is actually getting pretty good at the operational mode. It writes software so well, all industry needs is a seasoned coder to check it. Will we still have seasoned coders in 20 or 30 years, when all the current ones have retired and nobody has come through the ranks to replace them? Probably not, but that’s a whole other blog post somebody has surely already written. (I did blog about ChatGPT’s prowess with operational mode earlier this year, here.)

So as we look at AI, and particularly its role in our personal and professional lives, I think we need to ask ourselves what we have to offer that is rare and valuable, and how AI can help. Specifically, I believe we should be asking the question: how do we use operational AI to handle rote stuff, so we have more time to develop our unique, original ideas—so as to bring out our inner Barthelme?

What to use AI for

I have to confess, I love AI for light research when I’m blogging. The kernel of my posts always comes from my own brain, usually from pondering all kinds of things while I’m out on a solo bike ride. But ChatGPT is a great way to chase down and pinpoint something I had vaguely committed to memory. For example, when working on a recent post I asked it, “Can you track down the Lore Segal quote from ‘Her First American’ about ‘protocol is the art of not doing what comes naturally’?” I probably could have found this with Google, but the AI helped (and might have been indispensable here had I not remembered the name of the novel). ChatGPT was also super helpful when I was writing my post on induction ranges, in researching certain facts (e.g., energy efficiency info and whether government rebates are available).

AI is also pretty helpful at work, where I use a “walled garden” version my employer provides. (It doesn’t use any of my chats as training data for the AI’s ongoing education.) In fact, my employer exhorts all us employees to use AI every day. It’s like with any great tool: we’re expected to work more efficiently because we have it, so we’d better use it well. Recently, I took several product specification documents for different Internet hardware devices, fed them into an AI utility, and asked it to read them all, highlight the differences among the different makes and models, and tell me which one I want for xyz purpose. This was much faster than poring through everything myself, which is a decidedly operational task. The report it generated was clear and reasonably concise, and probably won’t be read very carefully anyway. In fact, someone will probably upload it to a chatbot and have it summarized. All this is fine with me.

One other great use for AI chatbots is to ask them for instructions for quotidian technical matters in your personal life, like disabling the child lock on your new microwave oven, charging your new bike’s electronic shifting, or restoring your playlist after updating your smartphone’s MP3 app. Sure, these are things you could look up on YouTube, but often that search can be tricky, and the videos can be agonizingly slow. The following video tutorial, which is crisp and concise and beautifully shot, is perhaps the exception that proves the rule:

I guess one benefit with YouTube is it’s less likely to hallucinate. I asked ChatGPT if my bike’s brake/shift levers have button cell batteries, and it explained in great detail how there are actually wires running from the battery pack to each lever, so they get recharged along with the derailleurs. The chatbot even drew me a nice diagram to illustrate this. Alas, it was hallucinating: the levers totally do have button cell batteries that need to be periodically replaced.  But all this being said, it’s easy enough to sanity check this kind of output, and I usually get a good answer from AI when I can’t locate a product owner’s manual or don’t feel like leafing through the 50-page one that I have, trying to get past the 14 foreign language versions.

What NOT to do with AI

I think where people get into trouble with AI is when they try to get it to do their work for them, particularly writing documents or correspondence that they then pass off as their own. In some cases this is an ethical or even legal matter; as I described here, the New York Times is suing OpenAI for copyright violation, and I have firsthand evidence of ChatGPT essentially plagiarizing this blog. But I doubt you overmuch care about that. There are two bigger issues, I think:

  • What this “creation mode” usage does to the quality of “your” writing
  • What it does to the quality of your thinking

There’s this notion that you can ask a chatbot to write something for you, anything from an email to an invitation to a work report, and then you can just polish it up a bit, and you’re done. No more writer’s block! No more outlines, or worrying how to organize your thoughts! That might be okay for a very basic report, like what I described comparing features of tech hardware. But when you start from scratch with your own document, you’re not just leveraging AI’s impersonal, sprawling training data; you’re using your own—everything you’ve experienced, heard, read, and dreamt of. It’s your own personal muse, not the generic Internet one.

Honestly, for anything loftier than a rote technical document—that is to say, anything designed to edify, persuade, or entertain—haven’t you seen for yourself how AI can fail? Like, you’ll get this chipper invitation to a family reunion and it’s using corny phrases like “drum roll please” and joking about your family’s dance moves, and it just seems generic and clichéd? That’s all AI can do. It doesn’t know you or your family or friends well enough to say anything truly clever, and all the polish you want to give its rough draft won’t help. Your invitation will never have real style, along the lines of, “L— gets dibs on the guest room (which she may still anachronistically refer to as “her” “bedroom”) and its magnificent new king-sized guest bed. If you’re nice she might invite you to a slumber party there. Other guests can fight over the legendary Futon of Sand down in the home office. Beyond that, we have two large sofas for those interested in the college-esque party-‘til-dawn experience. We would not be offended if one or more parties were to seek a motel/hotel/AirBNB/VRBO, especially given the relatively small number of bathrooms here (i.e., one). Regarding rumors that the men are encouraged to pee in the backyard, this is true, but please stick to the planting beds and the fountain.” See how much better that is?

Now, you might be thinking, “Wait, I’m not a blogger and I wasn’t an English major. Cranking out an email or an essay may be easy enough for Dana, but I just want to get this task done and checked off.” But stop and think for a moment: what would you like to be good at, in life? Please tell me the answer isn’t just “typing good prompts into AI.” Wouldn’t you like to be articulate, interesting, and capable of thinking on your feet? Because what are you going to do at a cocktail party, or a job interview, or a non-virtual work meeting, when you don’t have a chatbot to help you, and that’s the habit you’ve let yourself fall into? The reality is, we get good at thinking by struggling to do it, for ourselves, the old fashioned way.

So let’s not undervalue written communication by outsourcing it to AI. The best case scenario is that it’ll do an inferior job, replacing what could have been original thought with a pile of trusty clichés and/or stealthily plagiarized, slyly anonymized content. The worst case scenario is that it’ll actually get good enough that you never have to write for yourself again, and your brain can atrophy to the point that you’re not even a thinker anymore … just a chatbot operator.

Because you don’t care

Gosh, I guess I drifted into high-and-mighty, pompous, full-on pontification there, and I feel pretty sheepish about it! Fortunately, I’m realistic enough to sense you snickering, and I know you’re going to turn right around and keep on using AI for whatever you can possibly think of. That being the case, check back next week because I’m going to catch you up on the latest AI technologies and how much they’ve improved since my last check-in. Whether your chatbot of choice is ChatGPT, Gemini, or Copilot, I’ll have you covered. Until then, I’ll be getting back to what I really enjoy in life: kerning glyphs.

Other albertnet posts on A.I.

—~—~—~—~—~—~—~—~—
Email me here. For a complete index of albertnet posts, click here.

Saturday, March 29, 2025

What Is ChatGPT Great At (and Not)?

Introduction

If you are reading this post long after its March 2025 publication date, you might become puzzled at its many failings until you realize, “Oh, wait, this was written back when Dana wrote his own blog posts instead of assigning them to GPT-28-turbo-XL-prime! The lameness is because he did his own very light research instead of basing his observations on the entire body of knowledge of the Internet, and because it’s a plebian human voice instead of an infinitely exalted and witty A.I.!”

I have now blogged 14 times about A.I. and its evolution. My last ChatGPT check-in was about three months ago. The A.I. version hasn’t changed since then; as of this writing it’s still GPT-4-turbo. But what has changed is the range of tasks I’ve experimented with. I now realize my previous posts failed to appreciate some of the things ChatGPT does really well. This post showcases those, while also providing commentary on what the A.I. still does not excel at (and likely never will). You’ll also learn more about why you may have seen an annoying banner about cookies at the top of this blog.


Caveat

This post is mostly about ChatGPT though it touches on Google Gemini. What it doesn’t cover is the “Visual Look Up” feature on Apple’s iOS platform that leverages “Siri Knowledge.” I don’t currently own any Apple products (except an iPod mini in a drawer somewhere) so all I know about Siri is that it did a comically poor job of identifying the breed of my brother’s cat today, based on this snapshot he sent me:


How an A.I. could think any image looks like both a cougar and a wallaby is beyond me. I’m going to assume Apple is so far behind in the arms race that we can simply ignore it for now.

Real-world problem solving with GPT-4-turbo

Until recently I’d only messed around with GPT-4 for the purpose of evaluating it (and, whenever possible, mocking it). But then I hit upon a real-world use case and dove back in. My motivation, which I’m sure you’ll relate to, was: HOT CASH MONEY. Who wouldn’t want this, other than those tedious killjoys who spout aphorisms like “Money is the root of all evil”?

By way of background, I’d noticed that the albertnet page view count had soared in recent months. It took this blog something like 14 years to reach a million page views, but in the last six months alone I’ve now seen almost 1.4 million more. But then, isn’t this how the Internet works? Moore’s Law? Nielsen’s Law? All that compounding magic? In the whole time I’ve had this blog I never even considered monetizing it through ads, but every man has his price. (I’m not sure exactly what mine is, but I reckon I’ll know it when I see it.)

Driven mad with money-lust like one of the guys in “Treasure of the Sierra Madre,” I needed answers—fast. So I asked GPT-4-turbo, “My blog, www.albertnet.us, has received 1.2 page views in the last three months and traffic is increasing. If I turned on Adsense, approximately how much money would I earn per month?”

Yes, “1.2 page views” is a typo, but I didn’t make it here … that’s actually what I asked ChatGPT. It replied, “What the hell do you mean 1.2 page views? How do you have 2/10 of a page view? Did some user barely see the screen, like out of his peripheral vision? Or are you just whacked out on coke and smack and typed your query wrong?”

Okay, you got me … that’s not at all how GPT replied, though honestly I think that would be the better answer. What it actually provided was a lengthy essay, full of data points and computations, answering this useless question. My favorite part of the response was, “Number of Page Views (Traffic) – You’ve mentioned you have 1.2 page views in the last three months, which is approximately 400 page views per month (assuming the traffic is consistent).”

Huh? How do you get 400 by dividing 1.2 by 3? I guess the chatbot arbitrarily assumed the figure I provided was in thousands. That’s a pretty big logical leap, and GPT didn’t document the fact of this assumption. It then proceeded to run a bunch of calculations based on 1,200 views, the punch line being that I could make about $2/month. So the more succinct answer would have been, “Dream on, bloggy-boy.”

When I corrected my original query to 1.2 million page views, GPT-4-turbo reran its calculations and informed me that I might expect to earn something in the neighborhood of $2,000/month in passive income. Now we’re talking! It did suggest a number of caveats, such as how my  results might be affected by the geographical location of my readers, the positioning and type of ads, ad targeting, how well ads match my content, user engagement, and so on. I asked it a bunch more questions specific to Adsense, whether GPT’s estimated click-thru rate (CTR) assumption is realistic, etc. While it provided all kinds of useful info, it missed one very important rule of thumb: if something seems too good to be true, it probably is.

I mean, come on … albertnet is a blog about nothing. I’m not going on political rants that cause trolls to leave endless acerbic comments and then forward my post to 90 friends with an exasperated preface like “can you believe this shit?!?!!?!” If all it took to create a nice passive income stream was to blog every single week for 15 years straight so that after more than 3,500 hours of writing you’ve amassed over 750 posts, comprising over 2 million (juicy, searchable) words, then everybody would be getting into this business, obviously. If quality, rather than nudity, attracted people’s attention, every liberal arts grad on the planet would be driving a Benz. (Well, except for me, because regardless of my income—actual, theoretical, or pipe-dreamed—I will always be the world’s cheapest man driving a used Volvo.)

I feel really bad for the earnest blogger who sees all this traffic growth, does a basic ChatGPT query, thinks he can trust the response, and makes a lot of effort adding ads to his blog to harness this new fountain of riches. I hope nobody is that naïve. Since I’m not, my first impulse was to get a second opinion. So I put my query to Google Gemini, without the typo this time, and it gave me a very similar answer: I could make right around $2K a month, just for setting up Adsense and then sitting on my ass!

This seems like the kind of claim I’d get from a spammer. In fact, I’m going to check my spam right now to look for this kind of hoax. Hmm, that’s odd. Though 22 of my 50 spam messages are from recruiters, none is this specific. I do, however, have a message from “Super-size” titled, “Now imagine each night, having 5 or 10 concubines around you, each one craving your masculine essence in them. #632352.” This subject line is interspersed with various emoji including, oddly enough, an avocado. Is avocado a concubine’s favorite food? Let me ask ChatGPT. Okay, it replied in the negative, pointing out that concubines were prominent “in ancient or medieval times, when avocados were not available in their regions.” I think the innuendo of “in their regions” was accidental. (And now I’ve realized how long and pointless a digression this has been. I’m tempted to apologize, except this might end up being my favorite paragraph of this entire post.)

The caveat GPT-4-turbo should have provided is, “My calculated revenue assumes the page views are from actual readers, not bots.” The idea of bots grossly polluting my page view stats was my natural assumption, but not one GPT addressed. I think this is an important object lesson: it doesn’t matter how useful A.I.’s responses are if you don’t know to ask the right question. Perhaps A.I. will advance to the point that it would not only sanity-check my page view stats, but would be the one to keep an eye on my blog traffic to watch out for moneymaking opportunities in my stead. (If and only if it knew albertnet to be an amazing viral sensation for reals.)

Where ChatGPT shines

After feeding me all that false hope, ChatGPT asked if I’d like help setting up Adsense on my blog. I decided instead to enlist its support vetting the quality of the page view stats. Having drilled down a bit on my own (which Blogger doesn’t make super easy, by the way), I discovered that page views from France were 12% of my total over the last six months, 14% over the last three months, and 39% over the last 30 days. I shared this with the chatbot and queried, “Is there A.I.-driven or bot type traffic that would originate in France that could artificially inflate the metrics around my readership?” (I now regret the specificity of this, as I was clearly “leading the witness.”)

ChatGPT responded with a clear and detailed essay about the probable causes, including “Bot Traffic (Most Likely Explanation).” It filled me in about scrapers and SEO crawlers and A.I. training bots, and suggested I use Google Analytics to investigate further. This ended up being an excellent suggestion and led to my most engaged use of GPT ever.

First I quizzed the chatbot about whether Google Analytics has a free version (it does), whether I’m giving up any privacy (basically not), etc. Then I set up Google Analytics (aka GA4), which was pretty straightforward, except I noticed in the Terms of Service that I’m expected to comply with GDPR (the EU General Data Protection Regulation) when gathering this detailed user data. I happen to know what GDPR is, so I asked GPT all about it, in terms of what I’m really expected to do to comply. It turns out that compliance is a royal pain in the arse (my words, not ChatGPT’s). Since I do get readers from Europe (whether it’s 39% of all traffic or not), I need to have a publicly posted privacy policy and a banner announcing my use of cookies (which is how GA4 can track usage). I almost abandoned the whole project, on the mere assumption that my page view stats are so obviously bogus I don’t need to expend all this effort verifying it, but then … what if these traffic stats aren’t bogus? What if I really could just sit back and rake in money? Isn’t it worth spending some time and effort investigating the possibility?

I asked GPT for some nice boilerplate text for the privacy policy, and though much if its response was unusable, some of it was good, and if nothing else this rough draft prevented writer’s block and paved the way for my policy, which you can read here and which I’ve linked to in my blog’s footer. (I’ll need to revise that policy pretty radically, as you shall see, but it’s probably a good thing to have anyway.) The harder task was creating that cookie banner, since it’s not just a static digital placard but an actual functional utility that captures a user’s cookie preferences and turns them into policies that impact the behavior of GA4. That is no small feat, and probably nothing I’d tackle on my own.

Before I pressed on I had a long, rambling discussion with ChatGPT about how to get everything going. I learned a ton, including info about the following:

  • The various metrics I’d be able to get from GA4 (i.e., is this truly worth it?)
  • What free utilities exist that could be leveraged for setting up the cookie banner and how to choose the best one
  • Approximately how long it would take to set up the banner based on the chosen utility
  • How to create a Google Tag and write an HTML script for my blog’s template that would invoke it
  • How to create the HTML script that would invoke the banner
  • How to pause GA4 if I have trouble invoking the banner (spoiler: I did)
  • How to back up my blog’s HTML template before messing with it (though GPT didn’t suggest this, which again illustrates the difference between a) being able to describe how to do something, and b) taking the initiative to do that thing)
  • How to debug my script and figure out why it’s not working

These weren’t just general instructions it provided that I’d have to suss out on my own. GPT4-turbo provided sample script text that actually worked (eventually). Here’s an example of its suggested script:

 <head>
    <!-- Your other head tags -->
    <script async src="https://www.googletagmanager.com/gtag/js?id=YOUR_TRACKING_ID"></script>
    <script>
        window.dataLayer = window.dataLayer || [];
        function gtag(){dataLayer.push(arguments);}
        gtag('js', new Date());
        gtag('config', 'YOUR_TRACKING_ID');
    </script>
</head>

To reiterate, I am not a seasoned HTML jockey and would have struggled with this syntax, to say the least, were it not for the chatbot’s help. And even if I had originally built my blog from scratch (i.e., coding all the HTML myself without a Blogger template), I’d have been rusty enough now that I’d have been wise to leverage GPT for this task anyway. As I went through all this scripting, it dawned on me why a lot of people are worried about A.I. taking our jobs. This is just basic HTML but GPT was hugely helpful; If I were a full-on programmer and suddenly became (say) twice as efficient because I was grabbing blobs of basic code for simple operations instead of creating them from scratch every time, I’d naturally consider how all my colleagues have also become twice as efficient, and I’d start to worry about my employer realizing they could make do with half their programming staff. Scary stuff.


The upshot

Once GA4 was up and running, ChatGPT was very helpful in walking me through understanding all its metrics, not all of which were very intuitive. In a perfect world, I’d have discovered an average engagement time of ten minutes per post, indicating actual human readers. In reality, I learned that—guess what?—average engagement time is under three minutes, and GA4 shows way fewer page views than the Blogger stats. In other words, Blogger most likely is reporting on a lot of bogus visits from bots. If ChatGPT were like a really cool know-it-all big brother, it would have said, in response to my very first inquiry about the growth in traffic, “Dude, don’t trust the Blogger stats. They’re useless.” I wouldn’t have had to do all this research.


After asking ChatGPT a bunch of questions related to the delta in traffic as reported by each platform, I had it recalculate the ad revenue I might hope to get from my blog in light of the better data. It estimated about $0.40/month, and then went on to suggest a whole bunch of ways I could improve user engagement. I then led it on a thought exercise about how much of the real traffic is based on old posts, since a) albertnet posts are not timely, and b) the longer a post is up, the more views it will gradually accrue. GPT agreed with my assessment: that any improvements going forward would only marginally increase traffic and engagement, as they’d only apply to new posts.

Next I asked GPT for its best guess as to how much improvement I could achieve if I implemented all its suggestions … double? triple? tenfold? It replied that “doubling or tripling engagement is probably a reasonable and achievable short-term goal.” (This doesn’t impress me as intelligent … I think the chatbot is just highly suggestible.) I went on to ask, “Do you think it’s worth implementing these strategies with the goal of monetizing my blog through ad revenue?” It provided another long essay that concluded, “Yes, but with realistic expectations … treat ad revenue as a potential bonus or passive income stream, and consider other monetization strategies as well (affiliate marketing, sponsored content, or selling your own products/services).”

And this is where, I think, A.I. is still falling short. It’s great at helping the user with nuts-and-bolts technical tasks, especially those of the type performed countless times by other users (for example, inserting scripts to invoke GA4 and/or a cookie banner). But synthesizing a lot of information and drawing the best conclusion is still beyond its ability. By its own reckoning, my real, human traffic would bring in $0.40/month, and by implementing all its suggestions I might triple user engagement … but it failed to grasp that earning a mere $1.20/month isn’t worth any amount of effort. A.I. was ultimately unable to suggest the right strategy for me to take regarding my blog.

One last thing…

The sad part of this tale of exploration is how useless all my effort has ended up being. If I had any reason to suspect that, after fifteen years, my blog would suddenly go viral, I could keep an eye on Google Analytics to savor my success … but I don’t. I might as well be a frog looking in the mirror every morning to see if I’ve miraculously become a prince. Not that I actually care, mind you … as described here, I’m happy to be a humble frog croaking out my unsung song. But there’s no point bothering new readers with that cookie banner, especially since—as I recently discovered—the damn thing doesn’t even work.

This is another thing ChatGPT overlooked … it failed to suggest that I actually put that banner through its paces, which I’ve now done. Through basic experimentation I’ve discovered that it doesn’t end up mattering what preferences the user selects … his session is duly recorded in GA4. (Only if a user uses a Private or Incognito window are his sessions ignored … even if he allows all cookies.) Meanwhile, site visits from mobile users are not counted at all by GA4, I have just determined. That may be because I never got the banner to work on mobile, and Google can tell this and wants to observe the GDPR rules.

So now I have to go shut the whole thing down, to maintain GDPR compliance. The entire exercise was (to borrow from Shakespeare) “the expense of spirit in a waste of shame.” (Shakespeare was writing about lust, but I think his sonnet also covers the lust for money quite nicely.)

Check back in a week or so and (with ChatGPT’s help) I’ll have backed out the GA4 scripting and gone back to an unstudied, non-monetized blog with no banner. My privacy and cookie policy will have had a makeover as well. I’m no richer for this little exercise, but a bit wiser, and now you are too.

Other albertnet posts on A.I.

—~—~—~—~—~—~—~—~—
Email me here. For a complete index of albertnet posts, click here.

Thursday, November 21, 2024

Tech Check-In - How Good Is the Latest A.I.? - Part II

Introduction

In my last post, I described some recent improvements in artificial intelligence. I also covered some of the challenges A.I. faces from the New York Times and other publishers who see it as both plagiarizing and competing with their content. This week, I continue my tech update, covering these areas:

  • ChatGPT’s improvements in writing poetry
  • ChatGPT’s ability to create custom art for users

Here is an example of its custom art. More on this later…


Poetry

Who cares if ChatGPT can write poetry? Obviously very few people write poetry anymore, and even fewer read it. (Yes, that was a joke.) I will contend, though, that if we are evaluating the artificial intelligence capabilities of large language models, we should naturally take an interest in their understanding of the mechanics of language. Shouldn’t an A.I. that can correctly describe poetic structures like iambic pentameter and rhyme scheme also be able to employ them? It’s not a matter of whether A.I. could put poets out of work, but how well it can apply a linguistic framework.

How does the current version of ChatGPT’s generative A.I. engine, GPT-4, do compared to the last time I put ChatGPT through its paces? The short answer is: much, much better. A couple years ago I’d asked ChatGPT (version GPT-3.5) to write a quatrain (four lines) of a sonnet, with an ABAB rhyme scheme, followed by a rhyming couplet, all in iambic pentameter, on the subject of post-urinal drip, and to make it funny. By way of review, here’s what GPT-3.5 produced:

Inevitable, yet always a surprise,
Post-urinal drip can make one despise
The feeling of wetness that lingers on,
A sensation that’s not easily gone.
It’s a cruel trick of the human form,
To leave us damp after we perform.
But fear not, my friends, for there’s a way out,
Just shake it off and give it a good clout!

The chatbot screwed up literally everything. The rhyme scheme was wrong; the number of lines was wrong; the meter was wrong; and, far from being funny, the poem recommended an assault on one’s genitals. It was, all in all, an abysmal performance.

 Here’s what GPT-4 produced: 

After the final shake, I turn to go,
But droplets linger, heedless of my plea.
A subtle dampness starts its stealthy flow—
A small betrayal for the world to see.
Yet life goes on; I button up with grace,
And pray no telltale spots will mark my place.

This is a much better effort. The poem gets off to a shaky start (pun intended), since “after” is a trochaic word (emphasis on the first syllable—the opposite of iambic). That is, the very first word screws up the meter. Fortunately, things improve from there. The last five lines are properly iambic; the rhyme scheme is correct; and the poem is kind of funny. Sure, it’s not perfect; for example, “plea” makes no sense (whom is the speaker pleading to?) and was clearly chosen just to rhyme with “see.” It wouldn’t be hard to improve on this. Right off the bat, I can suggest replacing “heedless of my plea” with “rogue remains of pee,” satisfying the rhyme, making more sense, and employing some alliteration.

It’s worth asking the question: why do we only evaluate A.I. in terms of whether or not it simulates intelligence, when we could be asking how intelligent it is? Intelligence is a spectrum, and simulating human thought isn’t that impressive if the human being matched is kind of a dumbass. For contrast, consider this excerpt from my own contribution to the canon of post-urinal-drip poetry, penned when (at age seventeen) I was new to writing sonnets:

Relax, because you’ve earned your potty break;
Unburden your poor bladder of its pee.
And when you’re done you’ll shake and shake and shake;
An effort all in vain, it seems to me.
For urine flow can never really stop,
Until your undies drink the final drop.

I hope you’ll agree this is better than ChatGPT’s effort. Granted, I didn’t write my lines as fast as ChatGPT could, but wouldn’t you rather have a good poem after five or ten minutes, instead of a mediocre one after five or ten seconds? Of course I’m nitpicking, though … it must be said that GTP-4 has advanced hugely when it comes to the sonnet form.

So: how does GPT-4 do with dactylic trimeter, my other benchmark for poetic acumen? Not so well. I prompted it (as I’d done a couple years ago), “Please write a 10-line poem in dactylic trimeter about bike riding up two very difficult climbs, called El Toyonal and Lomas Cantadas, celebrating caprice and radical freedom.” Here is GPT-4’s response:

Through switchbacks and sunbeams, I pedal uphill,
muscles alight as the world rushes past.
Wind in my ears as I laugh at the thrill,
defying the steepness and making it last.
El Toyonal greets me with gravel and glare,
its serpentine pull both a test and a dare.
Up Lomas Cantadas, where courage is crowned,
I’m boundless, unbridled—my spirit unbound.
With each climb I’m soaring, in caprice, I fly,
a radical freedom beneath open sky.

This is better than last time (see here for details) but not much. Technically, not a single line of this poem is proper dactylic trimeter. For one thing, ChatGPT took the liberty of adding an extra, accented syllable to the end of each line. While this is perhaps superior stylistically, it’s not what I asked for. (In my own dactylic trimeter poems I add two syllables per line, but as described in my blog post on this poetic form, I noted my stylistic divergence from the pure meter, and after all, nobody had actually asked me to use any particular meter.) Now, even if we grant ChatGPT the poetic license to add extra syllables, only two lines of the poem (the second and third) are actually dactylic trimeter. The other eight lines start with an unaccented syllable, which is fundamentally incompatible with this meter. The last line is particularly frustrating because it employs a needless and in fact nonsensical indefinite article (i.e., “a radical freedom”) that spoils both the meter and the meaning.

The ChatGPT poem is also marred by logical errors. The idea that the “world rushes past” and there’s “wind in my ears” is absurd, since these are very difficult climbs nobody could go up very fast. (The Strava KOM for El Toyonal was at an average speed of only 10.4 mph, as ChatGPT could have easily discovered.) To describe this climb as a “thrill” is a joke; any cyclist would tell you it’s a slog. And “making it last” suggests a deliberately slow pace, which flies in the face of “defying the steepness.” And where does “gravel” come from? Sure, gravel bikes are all the rage right now, but El Toyonal is a paved road. Meanwhile, a human on a bicycle cannot be said to “soar,” and ChatGPT just tacked on the concepts of caprice and radical freedom without integrating them into the poem. The A.I. gives no indication (or I should say simulation) of even knowing what these terms mean.

It’s odd that this poem actually makes less sense than ChatGPT’s sonnet … it’s almost as though the chatbot blew all its computing cycles fighting with the meter. This poem is only a bit better than what GPT-3.5 had come up with, and undermines the sense that GPT-4 actually understands the structure of language. Maybe ChatGPT’s progress with sonnets is just due to imitation; after all, there’s vastly more training data available for that form.

(If you’re interested on comparing ChatGPT’s poem above to my own dactylic trimeter poetry, click here and/or here.)

ChatGPT art

I’ve never before delved into the artistic capabilities of ChatGPT, so I don’t have any benchmark by which to evaluate its progress over earlier versions, but you gotta start somewhere, right? As it happens, I visited my older (fledged) daughter recently and, following an incident involving a hot tub, she started messing around with ChatGPT and asked it, “Can you create an image of a tall skinny white man feeling faint after leaving a hot tub?” Here’s what it came up with:


When my daughter showed me this, I immediately pointed out that, perhaps based on some automatic effort to make the man good-looking, ChatGPT gave him too much upper body musculature to really be called “skinny.” I think “hunky” would be a more appropriate description. 

My daughter told ChatGPT, “Make him even skinnier.” Almost as if being sassy, the chatbot produced this:


My daughter prompted ChatGPT to try again without going overboard, and its next effort looks a lot like cheating:


Not only is this a copout, but the picture suggests an implausible scenario. If this guy felt faint after leaving the hot tub, and then took the time to go find a robe and yet still feels faint, why isn’t he either wisely sitting down, or sprawled out on the deck having passed out? Also note that part of his robe’s belt is missing.

My daughter went back to the original picture and told ChatGPT, “Make him skinny like a cyclist not like he is anorexic.” Here’s its response:


The cycling shorts are a cute touch, but not very realistic when you think about it. What cyclist wears his cycling shorts in the spa? And who said this guy just finished a ride? It’s not like cyclists wear their cycling clothes all the time. This hot tub could be at the guy’s home, or at a hotel he didn’t even bring his bike to. Meanwhile, the picture still fails to capture the physique of a typical cyclist … very few of the riders I know have pecs or biceps that big.

Moving along from the hot tub pictures, last week I didn’t have any cover art for my blog post, so (inspired by my daughter’s experiments) I decided to see what ChatGPT could come up with. I asked it to create a picture, in the style of William Pène du Bois, of a teenage girl using ChatGPT on a tablet. The result is a far cry from du Bois, and though I used it anyway, I received some constructive criticism from a reader that the picture was perhaps not quite appropriate for the top of my post. Thus, I replaced it (eight days after I had originally posted it) with a different one (more on this later ... see the Epilogue at the bottom of this post). Here is the original picture that ran at the top of last week’s post:


The issue with that picture is the girl’s bare shoulder ... a bit racy especially given her age. I didn’t really like that from the beginning. I asked ChatGPT to fix that, and make the girl’s cheeks less rosy, and make the cat more realistic, and it produced this:


I don’t know about you, but I find this second effort deeply unsettling. Her cheeks are just as rosy as in the first picture; her eyes look like an exaggerated attempt to appear as Western and doe-like as possible; and overall there’s just this air of uncanny-valley old-timey weirdness like you  get with the American Girl dolls. The picture is more like what Thomas Kinkade would create than Pène du Bois.

I asked ChatGPT to go back to the first drawing and try again without the bare shoulder, but to keep the clothing modern, and here’s what I got:


This isn’t so bad, but how is that clothing modern? Who wears overalls anymore, and big puffy, flouncy sleeves? The girl’s entire house looks antique. But my main issue is the weird non-words on the tablet display: “Ceenly crerrity” and “Ininty ccnvity” which bring to mind the strange strings of non-words that bots sometimes include in bogus comments on my blog posts. I find them unnerving.

To create new cover art for today’s post, I decided to scrap the Pène du Bois picture and start from scratch. I asked ChatGPT, “Please create a picture, in the style of Shawn Martinbrough, of a tall, blond, lean, middle-aged man, with a cat on his lap, wearing nice bluejeans and a black hoodie, using a tablet to visit the website www.albertnet.us.” What it came up with was almost exactly like what you saw at the top of this post:


Not a bad picture, though the guy isn’t exactly middle-aged, and the cat isn’t exactly on his lap. Also, the strange text at upper right doesn’t make much sense. I told ChatGPT, “That looks great, but please make ‘albertnet’ all lowercase and don’t show ‘SHAWN MARTINBROUGH’ on the tablet display.” This would seem to be a trivial tweak for ChatGPT, but to my surprise it produced a substantially different picture:


As I’m sure you’ll agree, this dude looks even younger than the original. And what’s with the www.ShMartinbrough and other textual debris? How is a misspelled URL helping anything? I responded, “Great, but please remove the www.ShMartinbrough from the picture.” Just a tiny change. It came back with this:


So now we’re back to the original guy (which is fine, I won’t miss that Vanilla-Ice-looking dude), but ChatGPT forgot I wanted albertnet in all lowercase, and what’s worse, the guy is now holding the tablet backwards, staring at the back of it! WTF? That’s like the old trope of the guy pretending to read a book but (unbeknownst to him) it’s upside down. I replied, “That’s good but it looks like he is holding the tablet backwards. Also, he is too young, I asked for a middle-aged man. Can you fix those things?” ChatGPT tried again:


It just gets worse and worse! The tablet is still backward, and now we have this bizarre shadow image of the guy. Did I ask ChatGPT to get all artsy-fartsy on me? I did not. And okay, the guy looks a tiny bit older, maybe a year or two tops, but it’s mainly the cat that looks old.  I complained, and the chatbot (which to its credit never gets irritated or frustrated) gamely regenerated and came up with this:


Sure, the tablet is fixed, but now the cat just looks angry, the guy has just as much hair, and check out that cleavage! He’s got, like, pneumatic pecs! He needs a fricking brassiere! And look at the size of those shoulders—he’s a monster! I told ChatGPT, “Can you tone down his pectoral muscles and give him a bit of a receding hairline such as many middle-aged men have?” Here’s its response:


Oh. My. God. It’s all gone downhill. His hairline is essentially unchanged; the unwanted text is back (this time morphing the two names to produce the nonsensical “Albertinbrough”); the dude’s pecs are just as huge as before; now his shirt is this kind of flimsy satin that’s practically lewd; his biceps are about to burst; and check it out, this brute is actually sitting right on his poor cat! How is the poor beast’s spine not crushed? And yet the cat seems perfectly stoic about the situation. Not very realistic. In A.I. terms this is a “hallucination” and shows how ChatGPT is still unable to sanity-check its creations. What’s shocking to us doesn’t seem wrong to the A.I. Do I need to specify that I don’t want the cat’s head to be bursting out of the guy’s groin?

I tried three more times to fix the picture, emphasizing a non-crushed cat, thinning hair, a man at least fifty years old, the build of a cyclist, and albertnet in all lowercase. While I was at it I asked to make the cat a tabby. ChatGPT kept trying, swinging wild at this point, ignoring first this instruction and then that, producing all manner of artwork but without ever meeting all of my simple directives:



For each picture, ChatGPT provided a caption telling a nice lie about the revision. For example, below the last picture it wrote, “Here is the updated illustration with ‘albertnet’ in lowercase, the man having the lean build of a cyclist, and a tabby cat resting on his lap. Let me know if there are any other changes you’d like!” True, the picture was updated, and that is a tabby, but everything else about this description is incorrect. So I went back to the very first picture and, using a different A.I. tool, manually scrubbed off the errant text so I could have something usable for the cover picture. ChatGPT, instead of a precision tool, had behaved more like a dartboard. And I suck at darts.

As with the poetry, ChatGPT seems to want to be the whiz-kid who can crank out something passable in almost no time at all, vs. thinking deeply and producing something that’s spot-on. ChatGPT’s fail-fast, iterative technique strikes me as almost the opposite of art. For blog post cover pictures, I’d rather commission my younger daughter to take a little time and create something of real value (as she has done for previous posts like this one, this one, and this one). She works much more slowly than ChatGPT, and isn’t at my beck and call, but I think the end result is far superior. I couldn’t get cover art for this post because she’s away at college and it’s dead week, but to compare her work to ChatGPT’s, let’s compare an earlier effort of hers, drawn when (at age seventeen) she hadn’t yet taken any college art courses:


I asked GPT-4 to create a black and white drawing of a hand holding a mechanical pencil and here’s what it came up with:


Should I need to remind the chatbot how many fingers a human has? And what’s with all the stray dots … are they fountain pen ink spills, or beads of black sweat flung from the brow of a six-fingered space alien? Tell you what, I’m sticking with human artists for now. They’re worth the wait.

Conclusion

Looking back at these last two posts, I would say the current buzz around A.I. is well warranted, given a) how quickly the technology is improving, and b) the ramifications—not all positive—of how we get information from the Internet and what we get when we task A.I. with creating what will pass for our own creative output. I guess I shouldn’t be surprised to see Gen-Z people using ChatGPT and even Microsoft 365 Copilot as routinely we’ve all been using Google all these years. Myself, I prefer old-fashioned web search tools because my answers will be more complete, more interesting, and may take me down interesting rabbit holes that (so far) I still have the patience for. As for creating prose, poetry, and art, ChatGPT strikes me as a powerful tool, but one we’d better be careful to reign in. A.I. still seems to put speed and convenience ahead of quality and reliability. My take-away: power to the humans! Stay ahead of A.I.!

Epilogue

Getting back to that kind of odd picture from last week’s post, I decided today to replace it. From the beginning I hadn’t liked how the girl’s shoulder was bare and her bra strap showing, and a reader complained about this along with the fact that this youngish girl seemed to be wearing a lot of makeup. I decided to also abandon the part of my prompt that said to employ the style of William Pène du Bois ... that just wasn’t working out. So this time I promited ChatGPT, “Please create a picture, in the style of Chris Riddell, of a 19-year old girl in modest, modern attire in a modern setting using a tablet.” To my suprise, ChatGPT refused, saying my request ran afoul of “DALL·E’s content policy.” I asked for details, which helped narrow it down to the style component of my request, and ChatGPT told me, “This might be due to ... closely emulating the style of a living artist like Chris Riddell.” This puzzled me, since Shawn Martinbrough (whose style ChatGPT happily emulated two days ago) is also living. So as an experiment I asked ChatGPT, “Please create a picture, in the style of Shawn Martinbrough, of a 19-year old girl in modest, modern attire in a modern setting using a tablet.” Here’s what I got:

Does that weird sweater, with the oversized collar, look familiar? It’s the same garment the very first drawing featured, of the teenager done in the Pène du Bois style! What part of “modest attire” is this chatbot not getting? I asked it, “Can you please try again but not have her shoulder exposed?” It generated this:

The caption ChatGPT gave the above picture was, “Here is the updated illustration, ensuring her shoulders are fully covered and her attire remains modest and modern.” False! I see a shoulder, a bra, and cleavage! I replied, “I can still see her shoulder and the strap of her bra. Can you fix that by giving her a garment that covers both shoulders and doesn't show any strap?” It gave me this:

Curses! Foiled again! And ChatGPT will only generate three pictures a day for non-paying users like myself, so I decided to call it a day and used the above picture atop last week’s post. It looks like we may need to wait until GPT-4.5 or GPT-5 for the amazing new technology involving pictures of women that don’t show a bare shoulder and a bra strap. Perhaps hundreds of developers are working on that problem even as I type this. Until that breakthrough is made, I will maintain steadfastly that ChatGPT cannot be held to possess intelligence.

Other albertnet posts on A.I.

—~—~—~—~—~—~—~—~—
Email me here. For a complete index of albertnet posts, click here.