Showing posts with label iambic pentameter. Show all posts
Showing posts with label iambic pentameter. Show all posts

Thursday, November 21, 2024

Tech Check-In - How Good Is the Latest A.I.? - Part II

Introduction

In my last post, I described some recent improvements in artificial intelligence. I also covered some of the challenges A.I. faces from the New York Times and other publishers who see it as both plagiarizing and competing with their content. This week, I continue my tech update, covering these areas:

  • ChatGPT’s improvements in writing poetry
  • ChatGPT’s ability to create custom art for users

Here is an example of its custom art. More on this later…


Poetry

Who cares if ChatGPT can write poetry? Obviously very few people write poetry anymore, and even fewer read it. (Yes, that was a joke.) I will contend, though, that if we are evaluating the artificial intelligence capabilities of large language models, we should naturally take an interest in their understanding of the mechanics of language. Shouldn’t an A.I. that can correctly describe poetic structures like iambic pentameter and rhyme scheme also be able to employ them? It’s not a matter of whether A.I. could put poets out of work, but how well it can apply a linguistic framework.

How does the current version of ChatGPT’s generative A.I. engine, GPT-4, do compared to the last time I put ChatGPT through its paces? The short answer is: much, much better. A couple years ago I’d asked ChatGPT (version GPT-3.5) to write a quatrain (four lines) of a sonnet, with an ABAB rhyme scheme, followed by a rhyming couplet, all in iambic pentameter, on the subject of post-urinal drip, and to make it funny. By way of review, here’s what GPT-3.5 produced:

Inevitable, yet always a surprise,
Post-urinal drip can make one despise
The feeling of wetness that lingers on,
A sensation that’s not easily gone.
It’s a cruel trick of the human form,
To leave us damp after we perform.
But fear not, my friends, for there’s a way out,
Just shake it off and give it a good clout!

The chatbot screwed up literally everything. The rhyme scheme was wrong; the number of lines was wrong; the meter was wrong; and, far from being funny, the poem recommended an assault on one’s genitals. It was, all in all, an abysmal performance.

 Here’s what GPT-4 produced: 

After the final shake, I turn to go,
But droplets linger, heedless of my plea.
A subtle dampness starts its stealthy flow—
A small betrayal for the world to see.
Yet life goes on; I button up with grace,
And pray no telltale spots will mark my place.

This is a much better effort. The poem gets off to a shaky start (pun intended), since “after” is a trochaic word (emphasis on the first syllable—the opposite of iambic). That is, the very first word screws up the meter. Fortunately, things improve from there. The last five lines are properly iambic; the rhyme scheme is correct; and the poem is kind of funny. Sure, it’s not perfect; for example, “plea” makes no sense (whom is the speaker pleading to?) and was clearly chosen just to rhyme with “see.” It wouldn’t be hard to improve on this. Right off the bat, I can suggest replacing “heedless of my plea” with “rogue remains of pee,” satisfying the rhyme, making more sense, and employing some alliteration.

It’s worth asking the question: why do we only evaluate A.I. in terms of whether or not it simulates intelligence, when we could be asking how intelligent it is? Intelligence is a spectrum, and simulating human thought isn’t that impressive if the human being matched is kind of a dumbass. For contrast, consider this excerpt from my own contribution to the canon of post-urinal-drip poetry, penned when (at age seventeen) I was new to writing sonnets:

Relax, because you’ve earned your potty break;
Unburden your poor bladder of its pee.
And when you’re done you’ll shake and shake and shake;
An effort all in vain, it seems to me.
For urine flow can never really stop,
Until your undies drink the final drop.

I hope you’ll agree this is better than ChatGPT’s effort. Granted, I didn’t write my lines as fast as ChatGPT could, but wouldn’t you rather have a good poem after five or ten minutes, instead of a mediocre one after five or ten seconds? Of course I’m nitpicking, though … it must be said that GTP-4 has advanced hugely when it comes to the sonnet form.

So: how does GPT-4 do with dactylic trimeter, my other benchmark for poetic acumen? Not so well. I prompted it (as I’d done a couple years ago), “Please write a 10-line poem in dactylic trimeter about bike riding up two very difficult climbs, called El Toyonal and Lomas Cantadas, celebrating caprice and radical freedom.” Here is GPT-4’s response:

Through switchbacks and sunbeams, I pedal uphill,
muscles alight as the world rushes past.
Wind in my ears as I laugh at the thrill,
defying the steepness and making it last.
El Toyonal greets me with gravel and glare,
its serpentine pull both a test and a dare.
Up Lomas Cantadas, where courage is crowned,
I’m boundless, unbridled—my spirit unbound.
With each climb I’m soaring, in caprice, I fly,
a radical freedom beneath open sky.

This is better than last time (see here for details) but not much. Technically, not a single line of this poem is proper dactylic trimeter. For one thing, ChatGPT took the liberty of adding an extra, accented syllable to the end of each line. While this is perhaps superior stylistically, it’s not what I asked for. (In my own dactylic trimeter poems I add two syllables per line, but as described in my blog post on this poetic form, I noted my stylistic divergence from the pure meter, and after all, nobody had actually asked me to use any particular meter.) Now, even if we grant ChatGPT the poetic license to add extra syllables, only two lines of the poem (the second and third) are actually dactylic trimeter. The other eight lines start with an unaccented syllable, which is fundamentally incompatible with this meter. The last line is particularly frustrating because it employs a needless and in fact nonsensical indefinite article (i.e., “a radical freedom”) that spoils both the meter and the meaning.

The ChatGPT poem is also marred by logical errors. The idea that the “world rushes past” and there’s “wind in my ears” is absurd, since these are very difficult climbs nobody could go up very fast. (The Strava KOM for El Toyonal was at an average speed of only 10.4 mph, as ChatGPT could have easily discovered.) To describe this climb as a “thrill” is a joke; any cyclist would tell you it’s a slog. And “making it last” suggests a deliberately slow pace, which flies in the face of “defying the steepness.” And where does “gravel” come from? Sure, gravel bikes are all the rage right now, but El Toyonal is a paved road. Meanwhile, a human on a bicycle cannot be said to “soar,” and ChatGPT just tacked on the concepts of caprice and radical freedom without integrating them into the poem. The A.I. gives no indication (or I should say simulation) of even knowing what these terms mean.

It’s odd that this poem actually makes less sense than ChatGPT’s sonnet … it’s almost as though the chatbot blew all its computing cycles fighting with the meter. This poem is only a bit better than what GPT-3.5 had come up with, and undermines the sense that GPT-4 actually understands the structure of language. Maybe ChatGPT’s progress with sonnets is just due to imitation; after all, there’s vastly more training data available for that form.

(If you’re interested on comparing ChatGPT’s poem above to my own dactylic trimeter poetry, click here and/or here.)

ChatGPT art

I’ve never before delved into the artistic capabilities of ChatGPT, so I don’t have any benchmark by which to evaluate its progress over earlier versions, but you gotta start somewhere, right? As it happens, I visited my older (fledged) daughter recently and, following an incident involving a hot tub, she started messing around with ChatGPT and asked it, “Can you create an image of a tall skinny white man feeling faint after leaving a hot tub?” Here’s what it came up with:


When my daughter showed me this, I immediately pointed out that, perhaps based on some automatic effort to make the man good-looking, ChatGPT gave him too much upper body musculature to really be called “skinny.” I think “hunky” would be a more appropriate description. 

My daughter told ChatGPT, “Make him even skinnier.” Almost as if being sassy, the chatbot produced this:


My daughter prompted ChatGPT to try again without going overboard, and its next effort looks a lot like cheating:


Not only is this a copout, but the picture suggests an implausible scenario. If this guy felt faint after leaving the hot tub, and then took the time to go find a robe and yet still feels faint, why isn’t he either wisely sitting down, or sprawled out on the deck having passed out? Also note that part of his robe’s belt is missing.

My daughter went back to the original picture and told ChatGPT, “Make him skinny like a cyclist not like he is anorexic.” Here’s its response:


The cycling shorts are a cute touch, but not very realistic when you think about it. What cyclist wears his cycling shorts in the spa? And who said this guy just finished a ride? It’s not like cyclists wear their cycling clothes all the time. This hot tub could be at the guy’s home, or at a hotel he didn’t even bring his bike to. Meanwhile, the picture still fails to capture the physique of a typical cyclist … very few of the riders I know have pecs or biceps that big.

Moving along from the hot tub pictures, last week I didn’t have any cover art for my blog post, so (inspired by my daughter’s experiments) I decided to see what ChatGPT could come up with. I asked it to create a picture, in the style of William Pène du Bois, of a teenage girl using ChatGPT on a tablet. The result is a far cry from du Bois, and though I used it anyway, I received some constructive criticism from a reader that the picture was perhaps not quite appropriate for the top of my post. Thus, I replaced it (eight days after I had originally posted it) with a different one (more on this later ... see the Epilogue at the bottom of this post). Here is the original picture that ran at the top of last week’s post:


The issue with that picture is the girl’s bare shoulder ... a bit racy especially given her age. I didn’t really like that from the beginning. I asked ChatGPT to fix that, and make the girl’s cheeks less rosy, and make the cat more realistic, and it produced this:


I don’t know about you, but I find this second effort deeply unsettling. Her cheeks are just as rosy as in the first picture; her eyes look like an exaggerated attempt to appear as Western and doe-like as possible; and overall there’s just this air of uncanny-valley old-timey weirdness like you  get with the American Girl dolls. The picture is more like what Thomas Kinkade would create than Pène du Bois.

I asked ChatGPT to go back to the first drawing and try again without the bare shoulder, but to keep the clothing modern, and here’s what I got:


This isn’t so bad, but how is that clothing modern? Who wears overalls anymore, and big puffy, flouncy sleeves? The girl’s entire house looks antique. But my main issue is the weird non-words on the tablet display: “Ceenly crerrity” and “Ininty ccnvity” which bring to mind the strange strings of non-words that bots sometimes include in bogus comments on my blog posts. I find them unnerving.

To create new cover art for today’s post, I decided to scrap the Pène du Bois picture and start from scratch. I asked ChatGPT, “Please create a picture, in the style of Shawn Martinbrough, of a tall, blond, lean, middle-aged man, with a cat on his lap, wearing nice bluejeans and a black hoodie, using a tablet to visit the website www.albertnet.us.” What it came up with was almost exactly like what you saw at the top of this post:


Not a bad picture, though the guy isn’t exactly middle-aged, and the cat isn’t exactly on his lap. Also, the strange text at upper right doesn’t make much sense. I told ChatGPT, “That looks great, but please make ‘albertnet’ all lowercase and don’t show ‘SHAWN MARTINBROUGH’ on the tablet display.” This would seem to be a trivial tweak for ChatGPT, but to my surprise it produced a substantially different picture:


As I’m sure you’ll agree, this dude looks even younger than the original. And what’s with the www.ShMartinbrough and other textual debris? How is a misspelled URL helping anything? I responded, “Great, but please remove the www.ShMartinbrough from the picture.” Just a tiny change. It came back with this:


So now we’re back to the original guy (which is fine, I won’t miss that Vanilla-Ice-looking dude), but ChatGPT forgot I wanted albertnet in all lowercase, and what’s worse, the guy is now holding the tablet backwards, staring at the back of it! WTF? That’s like the old trope of the guy pretending to read a book but (unbeknownst to him) it’s upside down. I replied, “That’s good but it looks like he is holding the tablet backwards. Also, he is too young, I asked for a middle-aged man. Can you fix those things?” ChatGPT tried again:


It just gets worse and worse! The tablet is still backward, and now we have this bizarre shadow image of the guy. Did I ask ChatGPT to get all artsy-fartsy on me? I did not. And okay, the guy looks a tiny bit older, maybe a year or two tops, but it’s mainly the cat that looks old.  I complained, and the chatbot (which to its credit never gets irritated or frustrated) gamely regenerated and came up with this:


Sure, the tablet is fixed, but now the cat just looks angry, the guy has just as much hair, and check out that cleavage! He’s got, like, pneumatic pecs! He needs a fricking brassiere! And look at the size of those shoulders—he’s a monster! I told ChatGPT, “Can you tone down his pectoral muscles and give him a bit of a receding hairline such as many middle-aged men have?” Here’s its response:


Oh. My. God. It’s all gone downhill. His hairline is essentially unchanged; the unwanted text is back (this time morphing the two names to produce the nonsensical “Albertinbrough”); the dude’s pecs are just as huge as before; now his shirt is this kind of flimsy satin that’s practically lewd; his biceps are about to burst; and check it out, this brute is actually sitting right on his poor cat! How is the poor beast’s spine not crushed? And yet the cat seems perfectly stoic about the situation. Not very realistic. In A.I. terms this is a “hallucination” and shows how ChatGPT is still unable to sanity-check its creations. What’s shocking to us doesn’t seem wrong to the A.I. Do I need to specify that I don’t want the cat’s head to be bursting out of the guy’s groin?

I tried three more times to fix the picture, emphasizing a non-crushed cat, thinning hair, a man at least fifty years old, the build of a cyclist, and albertnet in all lowercase. While I was at it I asked to make the cat a tabby. ChatGPT kept trying, swinging wild at this point, ignoring first this instruction and then that, producing all manner of artwork but without ever meeting all of my simple directives:



For each picture, ChatGPT provided a caption telling a nice lie about the revision. For example, below the last picture it wrote, “Here is the updated illustration with ‘albertnet’ in lowercase, the man having the lean build of a cyclist, and a tabby cat resting on his lap. Let me know if there are any other changes you’d like!” True, the picture was updated, and that is a tabby, but everything else about this description is incorrect. So I went back to the very first picture and, using a different A.I. tool, manually scrubbed off the errant text so I could have something usable for the cover picture. ChatGPT, instead of a precision tool, had behaved more like a dartboard. And I suck at darts.

As with the poetry, ChatGPT seems to want to be the whiz-kid who can crank out something passable in almost no time at all, vs. thinking deeply and producing something that’s spot-on. ChatGPT’s fail-fast, iterative technique strikes me as almost the opposite of art. For blog post cover pictures, I’d rather commission my younger daughter to take a little time and create something of real value (as she has done for previous posts like this one, this one, and this one). She works much more slowly than ChatGPT, and isn’t at my beck and call, but I think the end result is far superior. I couldn’t get cover art for this post because she’s away at college and it’s dead week, but to compare her work to ChatGPT’s, let’s compare an earlier effort of hers, drawn when (at age seventeen) she hadn’t yet taken any college art courses:


I asked GPT-4 to create a black and white drawing of a hand holding a mechanical pencil and here’s what it came up with:


Should I need to remind the chatbot how many fingers a human has? And what’s with all the stray dots … are they fountain pen ink spills, or beads of black sweat flung from the brow of a six-fingered space alien? Tell you what, I’m sticking with human artists for now. They’re worth the wait.

Conclusion

Looking back at these last two posts, I would say the current buzz around A.I. is well warranted, given a) how quickly the technology is improving, and b) the ramifications—not all positive—of how we get information from the Internet and what we get when we task A.I. with creating what will pass for our own creative output. I guess I shouldn’t be surprised to see Gen-Z people using ChatGPT and even Microsoft 365 Copilot as routinely we’ve all been using Google all these years. Myself, I prefer old-fashioned web search tools because my answers will be more complete, more interesting, and may take me down interesting rabbit holes that (so far) I still have the patience for. As for creating prose, poetry, and art, ChatGPT strikes me as a powerful tool, but one we’d better be careful to reign in. A.I. still seems to put speed and convenience ahead of quality and reliability. My take-away: power to the humans! Stay ahead of A.I.!

Epilogue

Getting back to that kind of odd picture from last week’s post, I decided today to replace it. From the beginning I hadn’t liked how the girl’s shoulder was bare and her bra strap showing, and a reader complained about this along with the fact that this youngish girl seemed to be wearing a lot of makeup. I decided to also abandon the part of my prompt that said to employ the style of William Pène du Bois ... that just wasn’t working out. So this time I promited ChatGPT, “Please create a picture, in the style of Chris Riddell, of a 19-year old girl in modest, modern attire in a modern setting using a tablet.” To my suprise, ChatGPT refused, saying my request ran afoul of “DALL·E’s content policy.” I asked for details, which helped narrow it down to the style component of my request, and ChatGPT told me, “This might be due to ... closely emulating the style of a living artist like Chris Riddell.” This puzzled me, since Shawn Martinbrough (whose style ChatGPT happily emulated two days ago) is also living. So as an experiment I asked ChatGPT, “Please create a picture, in the style of Shawn Martinbrough, of a 19-year old girl in modest, modern attire in a modern setting using a tablet.” Here’s what I got:

Does that weird sweater, with the oversized collar, look familiar? It’s the same garment the very first drawing featured, of the teenager done in the Pène du Bois style! What part of “modest attire” is this chatbot not getting? I asked it, “Can you please try again but not have her shoulder exposed?” It generated this:

The caption ChatGPT gave the above picture was, “Here is the updated illustration, ensuring her shoulders are fully covered and her attire remains modest and modern.” False! I see a shoulder, a bra, and cleavage! I replied, “I can still see her shoulder and the strap of her bra. Can you fix that by giving her a garment that covers both shoulders and doesn't show any strap?” It gave me this:

Curses! Foiled again! And ChatGPT will only generate three pictures a day for non-paying users like myself, so I decided to call it a day and used the above picture atop last week’s post. It looks like we may need to wait until GPT-4.5 or GPT-5 for the amazing new technology involving pictures of women that don’t show a bare shoulder and a bra strap. Perhaps hundreds of developers are working on that problem even as I type this. Until that breakthrough is made, I will maintain steadfastly that ChatGPT cannot be held to possess intelligence.

Other albertnet posts on A.I.

—~—~—~—~—~—~—~—~—
Email me here. For a complete index of albertnet posts, click here.

Wednesday, February 22, 2023

A.I. Smackdown — English Major vs. ChatGPT - Part 2

Introduction

In my last post, I considered the writing prowess of ChatGPT, the A.I. text generation platform powered by OpenAI’s GPT-3 engine. A professor quoted in Vice magazine said GPT-3 could get a B or B- on an MBA final exam, which I figured had to be an exaggeration. So, I put ChatGPT through its paces, having it write paragraphs in the style of a scholastic essay, a magazine article, and a blog post. In this essay, I’ll tackle a final writing category: poetry. At least when it comes to very logical matters such as rhyme and meter, A.I. ought to do really well … right? Well, let’s see how it does. (Hint: very poorly. And, I suppose I owe you a trigger warning: ChatGPT seems to think violence to one’s genitals is funny.)

But before I get to all that, I will address a pressing question: who cares about any of this? And why should we? I’ll also address a few follow-up questions from the friend who prompted my last post.

Who cares? And why should we?

In response to my previous post, my software maven friend suggested that A.I. could be used effectively for composition if the user (i.e., the person who’s tasked ChatGPT with responding to a query) has enough expertise to evaluate the response and wisely choose what to use from it, and what to discard. In this way, my friend suggests, “this version [of ChatGPT] could make someone who knows what they’re doing more productive.” He continued, “I wonder if for your next blog you might consider how you might use it? Would you be willing to use it as a first draft in responding to a low performing colleague who asks trivial questions?”

I have two responses to this: a practical one and an ideological one. On the practical side, I can’t imagine starting a work email, proposal, or report with ChatGPT because in my experience so far, most of what the A.I. does is apply window dressing and rhetorical flourishes to the ideas I feed it, along with vague assertions that aren’t backed up (e.g., “[Dura-Ace’s] sleek and understated style has been well-received by riders and industry experts alike”). ChatGPT builds repetitive, junior-high-grade essays that fill out the page but don’t add much value to the original prompt. If I were to obtain my rough drafts from ChatGPT, I would have to prune most of the text to end up with something reasonably concise. It would be faster just to write my own missive from scratch.

I realize this may not be true for everyone, and I’ll grant that I have developed, through decades of practice, uncommon facility with writing (having composed over 1.5 million words for albertnet alone). Nevertheless, the ability to quickly draft a work email or brief report is, I believe, a capability that any adult ought to have, just like being able to fry an egg, drive a car, or brew a good cup of coffee. To my mind, increased efficiency should be a matter of personal development, not outsourcing.

The ideological matter is more complicated. If we decide that producing a work document is the kind of hassle that should be dispatched with as little time and energy as possible, like submitting an expense report or making travel arrangements, we are diminishing the assumed value of that activity. As we prepare the next generation for the workforce, this sense of diminishment would trickle down to our schools. We would be sending students a message that writing is a job for A.I. and that the higher-value human thought lies elsewhere.

Having majored in English in college, I naturally bristle at this idea. I believe that reading and writing, more than so perhaps any other endeavor, teach us how to think. I’ve blogged before (e.g., here) about how strongly I disagree with American society’s obsession with STEM, as opposed to the traditional liberal arts that are all but dismissed in modern education. To those who promote STEM, I’d like to ask, what would you think about discontinuing most math classes in school, since we have calculators and spreadsheets to do that crap for us? Of course you wouldn’t support this, and neither do I. (I took a Calculus class in college just for the hell of it.) Studying math is good for your brain, even though most of the specific math skills you learn will never be used. Studying the craft of writing is also good for your brain, and using words well is a skill we can use every day of our professional and personal lives. Writing is hard, and takes time, sure. But when we strive to write well, we understand better, and we think more deeply.

Here is an interesting quotation from the American philosopher Alasdair Macintyre, quoted in the New Yorker, describing his misgivings about the Enlightenment:

It becomes impossible to settle moral questions or to enforce moral rules; the best we can do is agree to disagree. Such a world falls into the hands of managers and technocrats, who excel at the perfection of means but lack the tools with which to think deeply about ends.
I don’t have strong opinions about the Enlightenment, and MacIntyre was writing in the ‘80s, but I think that last sentence fits modern tech companies, particularly in the realm of A.I. development, rather well. What we see in tech is a lot of people asking “how.” How do we monetize Internet traffic, how do we create algorithms that increase this traffic, how do we increase platform dwell time, how do we make people forward more stuff, how do we make social media products more addictive. I don’t see a lot of people asking “why” or, crucially, “why not.” If we outsource writing to A.I., how will people learn to think more deeply and broadly, a capacity we may need now more than ever?

Okay, on to my main topic now. As you shall see, ChatGPT’s attempts at poetry show, in a particularly stark way, just how totally overrated this platform is today.

Shakespearean sonnet

When poetry is good, it can be transporting (think Shakespeare, Dickinson, and Eminem). But when it’s bad, it can be appalling (think Vogon poetry, the feared method of torture described in The Hitchhiker’s Guide to the Galaxy). Obviously this is largely a subjective matter, so in my analysis I’ll look mainly at what can be objectively measured: proper execution of meter and rhyme. This is highly logical stuff, almost mathematical, so it ought to be right up A.I.’s alley. I started with a classic form: the Shakespearean sonnet.

Now, in case you’re not familiar with that poetic form, A.I. is here to help. At least, it’s supposed to be. I asked ChatGPT, “What is a Shakespearean sonnet?” It replied:

A Shakespearean sonnet … consists of 14 lines and follows a specific rhyme scheme and meter... The structure of a Shakespearean sonnet is as follows:
  • Three quatrains (four-line stanzas) with a rhyme scheme of ABAB CDCD EFEF
  • A final couplet (two-line stanza) with a rhyme scheme of GG
  • Iambic pentameter meter, which means each line has ten syllables with a pattern of unstressed and stressed syllables

This description is pretty good (after I truncated it), though ChatGPT doesn’t get any special points for providing this, since any Google search could lead to you such an explanation, or a far better one such as this. What was missing from ChatGPT’s definition is a more specific explanation of which syllables are stressed vs. unstressed. (Get this backwards and you’ve got trochaic pentameter, which is totally different.)

So next I asked ChatGPT, “Please explain to me how to write a poem in iambic pentameter.” It came back with this:

[Misc. prelude omitted here for brevity.] An iambic foot is a set of two syllables where the first syllable is unstressed, and the second syllable is stressed. Here's an example of a line in iambic pentameter:

“To be or not to be, that is the question.”

[Misc. pointless repetition omitted.] Count the syllables in each line. Each line should have exactly ten syllables. Vary the stress pattern slightly to avoid a monotonous rhythm. [More BS omitted.]

Here, ChatGPT presents probably the most famous line from Shakespeare, rather than a really appropriate one. The line above is a poor example because if you count the syllables, you’ll see there are actually eleven, not the “exactly ten” ChatGPT correctly said there should be. In that line, Shakespeare broke the rules on purpose, to cause the reader and/or theater audience to stumble. This was to draw attention to the word “question,” which in the context of existence is the whole point of the soliloquy. Great writing, but not a great example if you’re really trying to understand iambic pentameter. “Shall I compare thee to a summer’s day” would be a far better choice.

Okay, sure, that’s a bit nitpicky, but the instruction to “vary the stress pattern slightly to avoid a monotonous rhythm” is completely wrong. If you vary the stress pattern, it’s no longer iambic pentameter! And the notion of monotonous rhythm? Total bullshit! Consider Shakespeare’s Sonnet 18, here. It is absolutely flawless iambic pentameter—not a single syllable out of place—and it’s brilliant throughout, not monotonous whatsoever.

Now, you might wonder what ChatGPT meant about “ABAB” (etc.) rhyme scheme, so I asked it, “What is an ABAB rhyme scheme?” It replied:

An ABAB rhyme scheme is a specific pattern of rhyming words in a poem or song. It indicates that the last word in the first line of the poem or stanza rhymes with the last word in the third line, while the last word in the second line rhymes with the last word in the fourth line, and so on. In this rhyme scheme, the letters “A” and “B” represent different rhyming sounds. For example, in the following stanza:

I walked along the beach in the bright sun,
And listened to the waves as they did run.
The sand was warm and soft beneath my feet,
And seagulls flew overhead in the heat.

The words “sun” and “run” rhyme, so they are represented by the letter “A”, while the words “feet” and “heat” rhyme, so they are represented by the letter “B”. This pattern continues throughout the poem or stanza, creating a consistent and pleasing rhyme scheme.

Did you catch that? How full of crap ChatGPT was just there? If not, let me make it a little clearer: let’s take those A and B labels and apply them to the four lines: you’d have A for “sun,” and A for “run,” and B for “feet,” and B for “heat.” It’s AABB. It’s not ABAB. It doesn’t match the first paragraph of ChatGPT’s response. And it’s not the rhyme scheme of a Shakespearean sonnet. (It’s not even the rhyme scheme of a Petrarchan sonnet, which is ABBA. It’s just a rhyme scheme ChatGPT kind of made up, which it uses for every poem I’ve seen it write.)

We have established that ChatGPT totally blew it at describing a Shakespearean sonnet. So how does it do when actually writing one? I told it, “Write a quatrain in iambic pentameter, with an ABAB rhyme scheme, followed by a rhyming couplet, also in iambic pentameter, about post-urinal drip. Try to make it funny.” Here’s its response.

Inevitable, yet always a surprise,
Post-urinal drip can make one despise
The feeling of wetness that lingers on,
A sensation that’s not easily gone.
It’s a cruel trick of the human form,
To leave us damp after we perform.
But fear not, my friends, for there’s a way out,
Just shake it off and give it a good clout!

Oh. My. God. Is this the same priggish ChatGPT that, as mentioned in my last post, balked at writing humorously about how bad a laxative drink tastes, claiming “I cannot provide humorous anecdotes about sensitive medical procedures” and “I do not generate inappropriate or insensitive content”? It’s talking about violence to one’s genitals! I had to look up “clout” just to make sure there isn’t some alternate meaning that wouldn’t be this awful. A.I. is advocating, albeit jocularly, giving one’s penis “a heavy blow with the hand or a hard object.” WTF?! That’s not funny at all!

Meanwhile, ChatGPT simply didn’t follow my instructions. I asked for a quatrain followed by a couplet. That’s six lines total; it gave me eight (i.e., two quatrains). The rhyme scheme is also not what I asked for. As for iambic pentameter, this poem is a train wreck. Certain brief phrases (e.g., “To leave us damp,” “Just shake it off”) are iambic, but not a single line of the poem achieves this meter consistently. The very first word, “inevitable,” is fundamentally incompatible with iambic pentameter. Chat GPT just wrote whatever the hell it felt like in a very loosey-goosey meter of its own design. If you’re not getting my point, read the above poem aloud and see how hung up you get. Then compare it to this example of an actual iambic pentameter quatrain and couplet, penned by a human:

Relax, because you’ve earned your potty break;
Unburden your poor bladder of its pee.
And when you’re done you’ll shake and shake and shake;
An effort all in vain, it seems to me.
     For urine flow can never really stop,
     Until your undies drink the final drop.

Note that this second example doesn’t resort to crude ideas like clouting your penis. And note how naturally the poem gallops along because it truly follows the meter. Also notice that it doesn’t have vague throwaway notions like “A sensation that’s not easily gone.” Then consider that I wrote those lines (full poem is here) when I was only seventeen. I guess someone accustomed to the crap poems printed on Hallmark cards might be impressed at ChatGPT’s effort, but I’m not. It’s almost cynical how casually it breaks all the rules, including that of decency.

Dactylic trimeter

A.I. learns through being fed “training data” from the Internet. In case its struggle with the sonnet was based on bad information, the Shakespearean form being a very popular one, I decided to give it another shot, with dactylic trimeter, a much rarer form I have come to enjoy. I asked ChatGPT, “Please explain how to write in dactylic trimeter.” It responded:

Dactylic trimeter … consists of three dactyls, or three stressed syllables, followed by an unstressed syllable, in each line… Count the syllables in each line. Each line should have exactly three dactyls, or three stressed syllables, followed by an unstressed syllable. This means the line should contain a total of seven syllables. 

The description above is completely wrong. A dactyl is not a stressed syllable. It is a three-syllable foot in which the first syllable is stressed and the second two are not. A line of dactylic trimeter has three feet, with three syllables per foot, for nine syllables total, not seven. And even if a line did have three stressed syllables followed by an unstressed syllable, that would be four, not seven. Is ChatGPT just stupid? If I have to start reading work emails written by ChatGPT, I think I’ll just blow my brains out.

Okay, let’s see how A.I. did at actually composing this type of poem. I told it, “Please write a 10-line poem in dactylic trimeter about bike riding up two very difficult climbs, called El Toyonal and Lomas Cantadas, celebrating caprice and radical freedom. It produced this:

Pedaling up Toyonal, heart pumping fast
My mind is clear, free from the world’s contrast
With caprice and freedom, I take each turn
Pushing through pain, for what I will earn
Onward I climb, up Lomas Cantadas
My legs scream out, but my spirit never fades
The wind in my hair, the world fades away
A moment of peace, that I long to stay
And at last I reach the summit high
A sense of triumph, that no words can deny

You can see right away what an awful poem this is. It’s what Thomas Kincade would produce if he were a writer instead of an “artist.” I mean, “free from the word’s contrast”—what does that even mean? And “the wind in my hair”—seriously? What is this, a Harlequin romance novel? Lomas is a 10% grade, and I’m wearing a helmet! But what particularly stands out is that again, ChatGPT didn’t follow my prompt whatsoever. In the entire poem, only two of the feet are proper dactylic trimeter (“pushing through” and the first three syllables of “Lomas Cantadas”), which is surely just luck. As it did with the sonnet, ChatGPT just wrote whatever the hell it felt like. So why does everybody praise ChatGPT so much? It sucks! (For a proper poem on this topic, with actual dactylic trimeter, click here.)

One more thing

Okay, I can almost hear you now: “Oh, this particular chatbot is just using GPT-3! The technology getting better all the time! All the glitches you’ve found will soon be fixed! The next version’s gonna be amazing!

Well, maybe GPT-4 (etc.) will get better at poetic meter, and maybe it’ll learn how to be more concise. But I could also imagine its errors getting propagated further. Remember, GTP-3 learned mostly from training on massive amounts of human output from across the Internet, and (as I learned from my software maven friend) has over 100 billion parameters allowing it in some sense to memorize an enormous portion of its training set. Over time, as more people outsource their writing to A.I., its errors could be added to the pile of training data, and thus reinforced. Meanwhile, the content may stray ever further from that created by humans. The growing body of text on the Internet may come to have less and less to do with us—that is, with creators who have a soul, and a conscience. It’s tempting to hope that somehow the works of great writers will one day be scored higher somehow, to help the A.I., but why would we expect this when politicians, the media, and academia are kicking liberal arts to the curb? Meanwhile, most social media platforms today seem to prize forwards and re-posts as the most valuable Internet currency, so if any scoring were to be applied to A.I.’s learning, it’s probably more likely to be whatever gets a rise out of people—i.e., trolling and other bombastic vitriol.

As ChatGPT and its ilk gain ever more traction, what passes for writing could become, to borrow a phrase from Nabokov, the “copulation of clichés.” (He was talking about pornography, but the metaphor holds here, too.) As the data set A.I. uses becomes more and more generic, while the tool gets used by more and more people seeking to avoid engagement with the craft of writing, most real insight and individuality might gradually vanish from written correspondence. O brave new world!

Other albertnet posts on A.I. 

—~—~—~—~—~—~—~—~—
Email me here. For a complete index of albertnet posts, click here.

Wednesday, July 8, 2020

English: I Think That That Language Is Screwy


Vlog

This post is available as a vlog. Put it on the big screen and gather the whole family! Or, fire it up on your phone, add earbuds, and pretend it’s a podcast!

Note: this post is not about Vladimir Nabokov or any of his works. It’s kind of inspired by his love of words, though, and there is some Russian language lore in here.


Introduction

If you’re a super-nerdy word buff like I am, this post should be right up your alley. If not, it’s an opportunity to silently mock me and feel relieved to be a normal person. Either way, read on!


The Poem

English: I Think That That Language Is Screwy

All languages are more or less complex.                            
To learn a second one is to subject                                   2
Yourself to much abuse. It does perplex                           
Me: so much stuff to get correct!

One language highly likely to abuse                      
The learner is our blasted English tongue.                     6
It’s almost like it’s tailored to confuse:
Of tangled threads it’s intricately strung.

Pronunciation rules do exist.
Exceptions, though, seem almost infinite.                   10
(To number them, much less produce a list?     
For such a project, I’m inadequate.)                    

As evidence, allow me to present                         
The heteronym, which proves without a doubt:         14
Our rules simply do not represent
An airtight way to sound our vowels out.

When I write “wind” what way have you to know           
If “moving air” is meant, or “reeling in”?                     18
Does “tear” suggest reaction to a woe,                
Or ripping something up? See? Evil twins!

Our consonants are dodgy, too. You see,
The word “refuse” can mean “debris,” although        22
The “s” can be pronounced just like a “z,”
In which case “turn down” is the sense, you know.

These heteronyms confound inflection, too.
Does “entrance” mean a way to get inside,                26
Or “put into a trance”? It’s up to you
To sort it out. It’s not well codified.
                                                                                                                           
What’s going on? There’s really no excuse         
For ambiguity to so abound.                                        30
Who authorized a system so abstruse
That letters make an arbitrary sound?

    The secret of our tongue, I think is that
    It does reward a grasp that’s intimate.                 34

Footnotes & commentary

Title: That That

No, it’s not a typo. The first “that” is a conjunction and the second is an adjective. Both are arguably necessary. More on this later.

Line 1: all languages

There has been some bickering among linguists and anthropologists about whether all languages are equally complex. As detailed here, there was a movement to assert this, perhaps driven by a desire to “reject any nationalist ideas about superiority of the languages of establishment.” I can’t see how anybody could objectively assess the complexity of this or that language, because we all have a native tongue that will make some languages easier for us to learn and others harder. Also, I don’t actually care.

Line 1: complex

Why is this word blue? And certain others? Hmmmm…

Line 3: abuse

Is “abuse” an overstatement? Perhaps, but I’ve seen and experienced some pretty disparaging behavior around the attempt to speak a foreign language. For example, a teenager I once knew, whose mom was from China, was visibly disgusted with her mom’s English and routinely said, “Learn how to talk—I can’t even understand you.” Myself, I struggled with French. A college instructor once told me, right before my oral final exam, “You should know: your pronunciation is terrible.” This didn’t exactly put me at ease. Then, halfway through the final, he stopped and said, “Before we go on, I just have to say, your accent is awful.” At the end he recapped how poor my performance had been throughout. Nice. It’s also worth noting that when I have  the classic anxiety dream of showing up to a final exam having never attended the class, the class is always French.

Line 4: so much stuff

There are so many possible errors you can commit with written communication. It occurred to me once, when reviewing my score on a college French quiz, that since the instructor is allowed to knock off half a point for any little goof, a poor student could end up with a negative score.

Incidentally, if you rightly recognized this poem as a sonnet (though it’s longer than the standard 14 lines), you may have noticed that this fourth line is missing a foot. That is, it has only four two-syllable feet instead of five. This is deliberate. With all five feet, the line sounded jarringly too long. I encountered  similar problem when I wrote my Ode to South Park. Its fourth line is technically correct (five iambic feet) but it doesn’t sound right. For eight years I’ve been considering fixing that.

Line 6: blasted English tongue

I am so glad I learned this language the easy way (i.e., from imitating my parents as a toddler). Its grammar is so much harder than that of Latin or Russian. I’m tempted to include French in this list of easier languages, but its insistence on arbitrarily assigning gender to an inanimate object, and then requiring articles and adjectives to match this gender, seems unnecessary and possibly malicious.

Line 8: tangled threads

I believe it’s pretty widely accepted that the difficulties inherent in English stem from its long and complicated history, starting with Germanic dialects that evolved over time by contact first with Vikings, then with conquering Normans, Bretons, and Frenchmen. Details here.

Line 10: exceptions … almost infinite

Other languages I’ve studied are so much more consistent than English. Sure, French has irregular verbs, but not all that many of them, and although my mouth cannot make the sounds this language requires, in principle its rules make sense, and the diacritics (accent marks etc.) are helpful. Russian is also very logical (in my experience, and according to a fluent speaker I consulted). As for Latin, my college class learned all the grammar in a single semester and spent the next term translating Cicero et al … that’s how consistent Latin is.

Line 11: number … list

I got sucked in by Heteronym web pages and ended up letting several hours of my life slip away. Existing lists on the Internet are either too short (like this one) or too long (like this one which lists 427 pairs). By “too long” I mean too generous. I don’t want to count theoretical heteronyms, like “luger” (the pistol—proper nouns shouldn’t count) or the word “as” when it’s pronounced “ass” to mean “a Roman coin.” Who’s ever heard of this coin?

Incidentally, my first attempt to list all the heteronyms lead me to the Heteronym Homepage, way back in 1996 when the Internet was pretty new to most people. In case you’re a youngster, I have to tell you that back then, websites weren’t the slick multimedia affairs we see today. They all tended to be like the Heteronym Homepage: some nerd at a university posting a little essay about heteronyms and inviting readers to contribute their own. (This was before blogs were a thing.) As you can see, I’m credited with three contributions (one of which I had to argue for and which the webmaster only begrudgingly posted, with a qualification). This was my first-ever Internet presence.

Line 14: heteronym

If the heteronym is my poster child for English being particularly hard, I suppose I should establish that it’s a mostly English phenomenon. Since there are something like six or seven thousand languages worldwide, demonstrating this would be like proving a negative, but among mainstream languages I believe this is true. More on this in the Appendix, if you’re interested.

Line 14: proves without a doubt

Of course the existence of the heteronym isn’t the only evidence we have that the English language is completely whacked. Consider the word “thought.” Why does this crazy assembly of letters, “o-u-g-h-t,” make an “ott” sound here, whereas in the word “drought” it makes an “out” sound? And how come removing the second “t” from “thought” changes the “o-u-g-h” from “ott” to “oh”?

When I lived in San Francisco, I always puzzled over the pronunciation of the street name “Gough.” It could rhyme with “bough,” “cough,” “dough,” “rough,” or “through.” The arrangement of the letters in English is so often useless. We have to learn so many pronunciations à la carte.

And what’s this business with “h” and how it affects other letters, like “s” and “c”? Why should the “c” in “ch,” which is a “k” or “s” sound, combine with “h” to produce a totally different sound than either of them makes alone? Makes no sense. You know how the Russians indicate the “ch” sound? They have a specific character for it: “ч.” For a “sh” sound they have the character “ш.” They even have a character, “ж,” for the “zh” sound we nonsensically suggest with a simple “s,” like in the word “pleasure.” And their “k” sound is indicated by a letter, “к,” that never makes an “s” sound like our two-timing “c.” (The French “c” also does double-duty, but at least when it makes an “s” sound they indicate this with a cedilla, i.e., “ç”.)

Line 18: reeling in

Yes, I acknowledge that “cranking up” (like a wind-up toy) might be a better way to convey this second sense of “wind” then “reeling in.” Call it poetic license: I had to set up the next rhyme (two lines later). I think I deserve some leeway, given how freaking hard this poem was to write. My original goal had been to pack a heteronym into every line, but that proved impossible (for me). If you count up the blue words you’ll see how far short I fell.

I’d also thought it’d be cool to use each sense of each heteronym, but you see I only managed that once.  I even wanted to rhyme two pairs so that one sense of each rhymed with the corresponding sense of the other. For example, I wanted to rhyme the noun “abuse” with noun “excuse” and the verb “abuse” with the verb “excuse.” Oh well. I don’t know why I thought I could do all this … probably I listen to too much Eminem. I have to remind myself: he’s a genius and I am not. (Plus, rhyming is easier for a rapper because he or she can warp pronunciations slightly to achieve the desired effect.)

Line 20: evil twins

This is not a reference to my brothers Bryan and Geoff, who are twins (though they were pretty evil as kids). It’s an allusion to the TV trope of a look-alike evil version of the hero. At least three episodes of the original “Star Trek” featured the evil twin concept; there was a “Magnum, P.I.” evil twin episode;  and if memory serves there were three evil twins in “Charlie’s Angels” once. I can’t think of a better way than “evil twin” to distill the heteronym concept.

Line 22: refuse

This word was a trap! Did you get it wrong the first time you read this stanza? Good, good. One characteristic of a proper sonnet is its strict adherence to iambic pentameter: you have to arrange your words carefully to naturally create the proper rhythm for the reader. As I’ve explained before in these pages, it wouldn’t do to  screw up your naturally iambic vs. trochaic words willy-nilly in a line of verse:

Right: Exquisite and expensive are her tastes
Wrong: Hot dogs are bad for foraging pit bulls
Right: The yuppie Zeitgeist sickens Uncle Ralph
Wrong: His blood pressure is getting acute now

In the first example above of incorrect verse, you’d have to put the stress on the third syllable of “foraging,” which just sounds wrong. And in the second wrong example you’d have to put the stress on the second syllable of “pressure” and the first syllable of “acute,” which is also unnatural.

It turns out this consistent, rhythmic inflection can often help the reader (subconsciously) recognize which sense of a heteronym is intended, if the word has more than one syllable. For example, in line 26, there’s not much context to suggest which meaning of “entrance” is the right one, but you probably got it right (EN-trance, a way to get inside) because the meter led you there.

So getting back to line 22, you probably read it “re-FUSE” the first time, until you got to “debris” and had to go back and revise your interpretation (perhaps subconsciously). If so, you probably noted a small snag in the rhythm of the poem. I employed this little trick to jar you a bit, to prep you for my point about inflection later on (line 25).

Line 26: entrance

Ever since I realized “entrance” was a heteronym, I’m unable to look at an “ENTRANCE” sign and think “EN-trance.” I always see “en-TRANCE,” as though there were a hypnotist on the other side of the door.


Line 31: abstruse

I deliberated about “obtuse” vs. “abstruse” here. Generally, I avoid using a fancy word where a simple word will do. But “obtuse” is just too much of a stretch for the meaning I’m looking for. It’s not just me: look how much one dictionary had to say about using “obtuse” to mean, well, abstruse:


Lunching with a colleague once, I described my entrée as insipid, and he said, “How can pasta be stupid?” I explained that insipid mainly means “lacking in flavor” even if the word is often used to mean “dull” or “generally lacking.” He refused to believe me, so I bet him $5. We stopped at a bookstore on the way back to the office to check a dictionary (this being in the pre-smartphone era). Easiest $5 I ever made. I should have tried to take $5 off my teenager today when she read this poem and asked why I didn’t use “obtuse” here.

Line 33: that

Why is this word in blue? Isn’t blue supposed to mean it’s a heteronym? Well, yes. I haven’t seen “that” listed on any of the heteronym web pages I’ve seen (though for that matter, none of them lists “misuse” either, which certainly belongs on the list). We know that for a pair of words to qualify as a heteronym pair, they have to be spelled the same, pronounced differently, and have a different meaning. For the first test, consider that at least three mainstream dictionaries (the three I’ve checked) show two different pronunciations for “that,” as shown here. 




That upside-down “e” character, ə, makes a sound something like “eh,” as in the words “about, item, edible,” etc. as shown here.

Now, you might say these are just alternative ways to pronounce the word, like we sometimes say “the” to rhyme with “thee” and sometimes to rhyme with “duh.” But I don’t think “thăt” vs. “thət” is arbitrary. When we use “that” as an adjective (to specify “the one singled out, implied, or understood”) we pronounce it “thăt” (rhymes with “hat”). But when we use “that” as a conjunction (to introduce a subordinate clause that “is joined to an adjective or a noun as a complement”), we say “thət” (rhymes with “pet”).

To test this (actually, to prove it, as I was already convinced), I wrote two sentences on Post-Its and took them around to my wife and kids to read aloud. The first read, “If you think that I’m going to put with that, you’re crazy.” All three read it as, “If you think thət I’m going to put up with thăt, you’re crazy.” No prompting was necessary: that’s how they naturally pronounced those words.

The next Post-It read, “I think that that that that man said is a lie.” All four read it, “I think thət thăt thət thăt man said is a lie.” By sounding the conjunctions with the “ə” sound, they were able to easily utter (and understand) the sentence, odd though it is. But when challenged to pronounce “that” to rhyme with “hat” in all four instances, they got tripped up. So: we have established a different pronunciation that tracks with the different meaning. Voilà! Heteronym!

(If it seems like I’m making an inordinately thorough case for “that” being a heteronym, it’s because my daughter drew me in to a spirited debate on the topic and got me all excited.)

Line 34: grasp that’s intimate

This really is the glory of English: by being hard to learn, it gives an unfair advantage to native speakers. All languages do this, of course, but to the extent that English is particularly difficult, the advantage is magnified. Moreover, English is a particularly good language to be utterly fluent in if you’re trying to a) be upwardly mobile in the global economy, or b) rest on your laurels. When I mention to people that I was an English major in college, I like to add, “It was a really easy major for me because I grew up speaking English at home.”

Another reason I contrived to end this poem with the word intimate? It rhymes perfectly with “thət” (as used in the previous line).


Postscript

I have just discovered another heteronym pair: poke (to prod) and poke (pronounced poe-KAY), the Hawaiian dish of marinated raw fish or seafood. This isn’t on any heteronym list I can find. Score!

Appendix

Is it truly the case that the heteronym is mainly an English thing? Well, I studied French for six years and never came across a pair in that language. According to this article, “French has relatively few heteronyms, and the ones they do have, they often put a gratuitous accent mark on one of the meanings to differentiate them. (For example, meaning ‘where’ and ou meaning ‘or.’ The accent grave makes no pronunciation difference for the letter ‘u’; it's added so that ‘where’ and  ‘or’ are not spelled the same.)” Wikipedia lists 21 French heteronyms, and points out that a heteronym pair is usually the case of a noun being spelled the same as one conjugation of a verb (particularly third person plural), for example “couvent” meaning either “convent” or “they brood” (as in eggs). Since the French don’t really pronounce “ent,” the verb form sounds quite different. But this is certainly more obscure than English heteronyms.

I never encountered heteronyms in Russian and can’t imagine them because every letter in that language so consistently makes a discrete sound. As for Italian, Wikipedia is ambiguous about heteronyms in that tongue. It states, “Italian spelling is largely unambiguous, with a few exceptions” but then lists 35 examples. I plugged half a dozen of these examples into Google Translate to try to hear the difference, and they all sound identical to me, certainly nowhere nearly as different as “tear” (the liquid) and “tear” (the verb). As for German heteronyms, Wikipedia says that language has “few.” Ditto Dutch. No other languages are mentioned, for what it’s worth.

Further reading


—~—~—~—~—~—~—~—~—
For a complete index of albertnet posts, click here.