Image Generation For Language Learning

Using AI image generation to make the visuals that language-learning materials need: a picture for each of thousands of vocabulary words, illustrations for a children's book, flashcards. A participant at the workshop raised it as a real possibility, since commissioning an artist for that much is infeasibile, especially for a small program.

Upsides

  • Scale and cost. Far cheaper and faster than an illustrator, which can make a picture-for-every-word feasible for a program that could not otherwise afford it.
  • Context for learners. A picture conveys a word's meaning directly, without routing through a dominant language.

Downsides

General-purpose models often lean on stereotypes, and prompt engineering is not always effective. Asked for "a Native American boy going to school," Dall-E returns a child in a feathered headdress:

imagegen-native-boy-no-prompt.png

"A Native American boy going to school," with no further instruction.

Telling it to use "casual school clothing, without traditional elements" turns the clothes a bit more casual but keeps the headdress anyway:

imagegen-native-boy-prompt-engineered.png

The same request, prompt-engineered to avoid traditional dress.

Using image generation for curriculum development has to be done carefully. Left unchecked, it can amount to misrepresentation and digital redface at scale. State-of-the-art models inherit the stereotypes in the datasets they learn from, so the headdress above is cannot be reliably prompt-engineered away. Also, community artists often do not want their work absorbed into these models, but leaving it out only makes the output more stereotyped. This is a question of ownership and consent with no easy answer.

Created · Updated
Supported By the National Science Foundation Award 2542375.