Essay
Seeing Past the Camera: The elision of illusion and the actual, revisited
A picture of the Pope in a white puffer jacket was taken for a photograph by a great many people in March 2023, and the mistake has a long history: Leonardo’s buckler, the Lumières’ train, the rendered spaces of games. This essay asks whether anything but the medium has changed. Its answer is that the process is the same, an audience learning to read a new kind of image for the signs of how it was made, while the generated image differs from the render in the respects the essay follows, coverage, accessibility, scale and quality, each a matter of degree. The doubt now runs the other way as well, since real images are taken for generated ones, and the question is asked of every image at once, including the photographs the generators were trained on. It has moved from how an image was made to where it came from.
The story of being fooled
On Friday 24 March 2023 an image of Pope Francis in a long white puffer jacket was posted to the Midjourney subreddit, spread on Twitter over the weekend, and by the Monday NBC News was reporting that the pictures “fool the internet” (Tolentino, 2023). An audience mistaking a new kind of image for the thing it depicts is an old story, and Vasari, writing in 1568, tells it of Leonardo’s buckler, painted with a monster that made Ser Piero start at first sight (Vasari, 1568/1991, p. 287). It is told of the Lumières’ train in 1896, although Bottomore found that panic was rare and that accounts of it were spread partly “from motives of publicity by some showmen” (Bottomore, 1999, p. 201), so that the story circulated whether or not the event had. It recurred with computer-generated images, in photographs of game spaces where a virtual camera was given the marks of a real one so that a rendered space would read as captured (Paterson, 2013), and the puffer jacket is one more instance of it rather than a new kind of event.
Each of these is a case of one fascination, with an image that can pass for what it shows, and the fascination recurs with each new medium rather than belonging to any one of them (Paterson, 2013, pp. 9, 12). The move from the train to the game has been made before, in Gunning’s terms of the attraction rather than the deception: “From train effect we move to a game effect. From the cinema of attractions we move to the cinemas of interactions” (Gurevitch, 2010). The generated image is read here as the next case in that sequence, a continuation of one process in which an audience meets a new kind of image and learns, over time, how to read it.
The question this essay asks is whether anything but the medium has changed. It argues that the process is the same, that the image differs from those before it, and that the differences that matter are of scale rather than of nature. To understand what has changed it is first necessary to set out how an audience learns to read a new image at all, and what the signs of a camera established. The essay then follows those signs to the point where they were separated from any camera and asks what a generator adds to that separation. The second half turns to the audience: what the question audiences ask has become, what the eye can still settle, where the doubt now runs, and what is being replaced while it does.
How an audience learns to read a new image
In tests run in 2012, participants were shown photographs of game spaces mixed with photographs of physical ones and asked which was which (Paterson, 2013, pp. 55–56). They read the images for the qualities of the render: harsh light against the diffused light of a cloudy day, a large depth of field where the photographs had a short one, the look of low resolution and an over-sharpened feel. One participant read the arrangement of the bushes instead, since game developers tend to space objects too evenly when they imitate randomness, and identified the origin of every image (p. 56). Both readings look past what the image shows to how it was made. Socialisation to a new kind of image consists in learning which signs point to which process.
The signs are read first for the process, and once an audience knows them they become a vocabulary a maker can use, so that “the photographic as a rich vocabulary of conventions and references lives on” whatever happens to the camera (Batchen, 2000, p. 109). Game images are a case of that vocabulary in use, since their audiences “are already socialized to perceive visual media through the lens of the camera” (Paterson, 2013, p. 27), and how the signs came to be trusted lies in how the camera makes its image: “for the first time an image of the world is formed automatically, without the creative intervention of man” (Bazin, 1960, p. 7). The absence of the hand of man is what allowed photographs to be read as indexical (Paterson, 2013, p. 35). The distinction that follows is between the constructed image and the captured one: in a painting every mark is intended, whereas in a photograph the photographer chooses the frame and everything within it that was in front of the lens is taken in, intended or not.
The constructed image and the captured one are also true in different ways. A formula attributed to Pierre Bonnard, and best known as the epigraph to Josipovici’s novel about him (Josipovici, 1986/1998), puts the painter’s way as many little lies told for the sake of one large truth. It describes the constructed image well, since no single mark is the subject, and the whole can still be a true depiction because a person directed each departure from appearance. The captured image makes its claim the other way round, through every part of the frame at once, since nothing in it was put there by a hand. Which of the two is in front of a viewer is therefore the first thing the signs are read for, a reading that has to be learned.
Because the reading is learned, it is set by when the viewer learned it. Gunning’s account of novelty has astonishment that “gradually gives way to an acceptance of the new technology as second nature” (Gunning, 2003, p. 40), and adds that this second nature “may be less stable than we think” (p. 39). The instability appears to be generational, since a viewer born into the habit never experienced the astonishment. A phone plays a shutter sound although nothing in it moves, so a viewer can meet the simulated sound before the mechanism it imitates and still read it as the sign of a photograph being taken, which leaves the sign pointing to a process the viewer may never have encountered.
The camera established less than the trust later placed in it would suggest. A photograph establishes that something was in front of the lens, and does not establish what that something was. Its truth is therefore the reverse of the painter’s, lying in the parts and not in the whole, since each thing in the frame was there and what the whole shows is left to the viewer. The signs of the camera answered one question, whether a lens had been involved, and the first images to carry those signs without one were not generated but computed.
When the apparatus was first removed
Manovich, writing in the mid-1990s, observed that computer graphics defines realism as the ability to simulate an object so that its image is indistinguishable from its photograph, and that what the field had achieved was therefore “not realism, but only photorealism” (Manovich, 1996). The first images to carry the signs of a lens without one were made for film, where the signs had to be put in on purpose. To composite computer-generated elements into Jurassic Park in 1993, “computer-generated images had to be degraded; their perfection had to be diluted to match the imperfection of film’s graininess” (Manovich, 1996). The apparatus was gone and its imperfections were added back, which is why Manovich could conclude that three-dimensional computer graphics “can also be thought of as digital, or synthetic, photography” (Manovich, 1996).
Game engines did the same in real time, since a game’s camera is a construct with no lens and “complex algorithms are used to mimic the physical camera with, for example, the addition of lens flares and depth of field” (Paterson, 2013, p. 27), together with the sway of a handheld walk. The same mark can therefore read two ways, depending on the apparatus the image is trying to pass for. In a photograph a flare is often a fault, the lens getting between the viewer and the scene, whereas in a render it is put there on purpose and is the sign that a lens was involved at all. The render was built to be read as captured, and a viewer who knows the signs of the camera can recognise what it is asking of them, whether or not they grant it.
The socialisation of audiences to this image is visible in what players did with it. Screenshots came to serve “much the same function as the photo in physical environments: pointing out events and occurrences, documenting a sight seen” (Poremba, 2007, p. 50), and large games now build a camera in: Red Dead Redemption 2 added a Photo Mode with “free-form camera movement” and filters (Rockstar Games, 2019b), and The Legend of Zelda: Tears of the Kingdom has an in-game camera and a compendium to fill with its pictures (Hagues, 2023). The practice has a community and a name, and one of its practitioners describes “exactly the same skills as all the photographers do… depth of field, silhouettes, exposure” (Spies & Urban, 2026). Bogost and Poremba note that games are “celebrated for their realism, but this is typically a reference to their visual verisimilitude rather than an association with something actual” (Bogost & Poremba, 2008, p. 12), which is Manovich’s photorealism seen from the player’s side.
The question this image put to an audience was image origin, whether a photograph showed a physical space or a virtual one, and in the tests described above it was answered correctly 88 per cent of the time (Paterson, 2013, p. 55). The two images participants found hardest were both interiors from the virtual space, an office and an airport (p. 55). The question could be settled by looking, imperfectly, and on the render’s signs. A generator produces the signs of a lens without one, as the game engine did, and the next part sets out how it differs from the game engine.
What generation does differently
The opening of The Shining is filmed from the air, and Kubrick’s account of it shows what such a shot once implied: he “hired Greg McGillivray, who is noted for his helicopter work, and he spent several weeks filming” (Kubrick, quoted in Ciment, 1980). For an audience in 1980 the shot implied a helicopter, a specialist, weeks of flying and a studio able to pay for them. The same shot cost little more than a lens once the first DJI Phantom with a camera built in was announced in October 2013 (DJI, 2013), and a viewer who has attributed aerial footage to drones for as long as they have watched anything receives the framing without the prestige. An image generator needs no aircraft at all, and it can produce the shot because its corpus contains the images of everyone who did fly.
The image generators this essay concerns are diffusion models, and a diffusion model is trained by adding noise to images until nothing of them is left and learning to reverse each step (Ho et al., 2020; Rombach et al., 2022), on a corpus of billions of captioned images gathered from the web, 5.85 billion pairs in one open set (Schuhmann et al., 2022). Given a prompt, it starts from noise and removes it until what remains is the kind of image the corpus made probable for those words, and no camera, virtual or otherwise, is pointed at anything. The lens flare and shallow focus the game engine had to add are here neither added nor absent, since they are in the corpus as properties of photographs in general rather than as decisions about this one. The marks of every other kind of image in the corpus are there on the same terms, brushwork and halftone and film grain among them. Steyerl calls the result a mean image: “statistical renderings, rather than images of actually existing objects”, which “replace likenesses with likelinesses” (Steyerl, 2023).
Four differences from the render follow from the corpus, each of them a matter of scale rather than of nature. The first is coverage, since a game engine imitated one apparatus, the camera, whereas a generator produces the signs of any apparatus and any aesthetic its corpus contains, and the prompt names which is wanted. The second is accessibility, and it has a precedent: the camera allowed “the rendering of naturalistic images by people who were unskilled in the art of painting” (Paterson, 2013, p. 12), rendering then demanded expertise again, and generation demands a sentence. It became a consumer medium inside one year, with Midjourney’s open beta announced on 13 July 2022, DALL-E 2’s beta widened on 20 July toward a million people on its waitlist and Stable Diffusion released on 22 August (Midjourney, 2022; Wiggers, 2022; Lopez, 2023). The third is scale, which is a matter of how the reading has to be done rather than of how the image is made. The Cottingley plates were examined at the Kodak offices by two of the company’s experts, reported on in writing by another expert, and enlarged on a lantern screen at Wakefield before a sceptical operator (Doyle, 1921, pp. 32–33, 53, 91–92), for two negatives. An examination of that kind is not available for each image at the rate images now arrive, so the reading has to be constant where it was once occasional. The fourth is quality, since the output is now good enough that the cues which would prompt the reading at all are going, as the next part’s measurements show.
Each of the four takes the process described earlier further than the render took it, without starting a new one. The signs of any apparatus can now be produced by anyone, in quantity, and well enough to pass, and under those conditions the question an audience puts to an image cannot stay the question it was.
Real or generated
The question has a community named for it, since a subreddit called r/RealOrAI asks of each image posted whether it is a photograph or a generation, and a year of its activity, more than ten thousand comments, has been studied for how its members decide (Elmas, 2026). The question is the one the 2012 tests asked, with the contrast term moved. It is no longer whether a photograph showed a physical space or a virtual one but whether the image was made by any apparatus at all, and the renders that were once the other side of the question now count as real. In April 2023 the question was put to an institution, when Boris Eldagsen entered a DALL-E 2 image in the Sony World Photography Awards, won the Creative category of the Open competition, and refused it: “AI images and photography should not compete with each other in an award like this. They are different entities. AI is not photography. Therefore I will not accept the award” (Eldagsen, 2023), and his proposed name for what he had made was promptography.
The studies published since 2022 measure how well the question is answered by looking, and on faces it appears to be settled, since Nightingale and Farid concluded that synthesis engines “have passed through the uncanny valley” (Nightingale & Farid, 2022), and Miller et al. found that “White AI faces are judged as human more often than actual human faces”, with the participants who made the most errors being the most confident (Miller et al., 2023). Scenes are less settled: a quiz of about 287,000 judgments scored 62 per cent overall and lowest on landscapes (Roca et al., 2025), and a study of landscapes and interiors found 29 per cent for one newer model, its participants describing the images as “too perfect” (Högemann et al., 2025), which is the phrase Manovich had used for computer graphics thirty years earlier. These figures are not comparable with the 88 per cent of 2012, which tested different images with different people on a different question, so the comparison that matters is in the cues rather than the scores. Kamali et al. sort the cues that give away a generated image into five categories, “anatomical, stylistic, functional, violations of physics, and sociocultural” (Kamali et al., 2024), and on this reading only the stylistic one concerns how the image is rendered, while the rest concern whether the scene could be coherent. The reading the bushes participant made in 2012, unusual then, has become the ordinary method.
It is also a method whose cues move with the models, since the study that found 63.7 per cent overall found 29 per cent on the newest model it tested (Högemann et al., 2025). Training helps within that limit, since technique tips improved detection where, in the study’s own title, “A Warning is Not Enough” (Huang & Hu, 2025), and five short interventions raised discernment “by up to 13 percentage points” (Geissler et al., 2026), a real gain, made of cues the next model may not leave in place. Golby checked the hands in March 2023 because, as he put it, “AI struggles with hands” (Golby, 2023), and by 2024 the hand is the kind of cue Kamali et al. file under anatomical (Kamali et al., 2024).
Two things can happen when an image passes as a photograph and is not one. Gunning’s early audiences were “sophisticated urban pleasure seekers” (Gunning, 1989/1995, pp. 116–117) and what they came for was “a pleasurable vacillation between belief and doubt” (pp. 116–117), and the puffer jacket was posted to a community that makes such images in order to admire what the machine had managed. When the same image crossed to a feed with no frame around it, a viewer took it for a photograph. Nothing in the pixels had changed, only the context, which moved the viewer from participant to victim. Two conditions are therefore needed for a viewer to be fooled rather than entertained, that the image passes and that the context fails to say, and the studies above measure only the first. Golby’s column appeared three days after the image and began “It happened to me” (Golby, 2023). He had pinched and zoomed, checked the hands, and posted it anyway, and he records that the tools “seems to have leapt forwards” since “last year”. Earlier media required “a slow and ongoing audiovisual socialization” (Paterson, 2013, p. 15); this one ran in public and within months.
The doubt also runs the other way, from the generated image to the genuine one. Chesney and Citron named the liar’s dividend, by which a public that knows fakes exist can be told that a genuine recording is one, and the dividend “flows, perversely, in proportion to success in educating the public about the dangers of deep fakes” (Chesney & Citron, 2019). Every real recording becomes deniable, and the better the audience is socialised the more deniable it is. The evidence on how well the tactic works is mixed, with one study finding the dividend paid and another finding it largely refused. Grohmann, Halle and Appel measured a false claim that a real video was a deepfake raising a politician’s perceived leadership (Grohmann et al., 2026). Schiff, Schiff and Bueno, across five experiments with more than 15,000 participants, found that such claims “prove effective against text-based reports but largely fail against video evidence” (Schiff et al., 2025). The two studies differ on the size of the effect rather than on whether it exists. A doubted photograph is not itself new, since digital manipulation raised the same doubt a generation ago and Batchen’s answer to it was quoted earlier, and what is new is that the doubt no longer attaches to one kind of image. If a real photograph can be taken for a generation, the doubt is not confined to generated images, and the next part follows where it goes.
Every image at once
The doubt reaches the images the generators were made from, since the corpus described earlier was gathered from the web. The aerial images a generator draws on when it produces the helicopter shot were made by people who flew, and once any such image can be generated, the ones that were flown for are open to the same question as the ones that were not. The camera in a pocket now adds to the doubt from its own side. Google’s Pixel 8 offers a feature called Best Take, which, as TechCrunch described it, “combines multiple group photos to create a version where anyone is not blinking or looking away”, so that a user can “choose expressions of different people from various shots and merge them into one” (Mehta, 2023). The result is a photograph of a moment that did not occur, made by the camera rather than against it, which is what an ordinary group picture may now be. The doubt therefore attaches to no new class of image but to the image as such.
The strongest case against that conclusion is made for one class in particular. Bátori argues that images generated from a user’s own photographs “possess photographic indexicality in their constituent parts… while lacking indexicality as wholes”, that they are comparable to photomontage, and that they belong to photography as a medium because the user keeps “meaningful human control” over conception, prompting, iteration and selection (Bátori, 2026). The scope matters, since the argument is made for images prompted with photographs rather than for generation from text, and within that scope the first half of his description is exact, since the parts are answerable to a photographer. The whole is answerable to nobody, which is the essay’s reply rather than his claim, and set beside the two kinds of truth described earlier it is the photograph’s truth without the painter’s. Bonnard’s small departures served one large truth because a person directed every mark, whereas the person Bátori describes directs the request and the choice among outputs, and the particulars inside the frame come from the corpus. Meaningful control over the prompt is real without being control over what the image shows.
Cottingley shows what provenance was doing all along, though not because the examinations failed. Doyle took the negatives to the Kodak offices, where two of the company’s experts “could find any evidence of superposition, or other trick” in neither, although they “would not undertake to say that these were preternatural” (Doyle, 1921, pp. 32–33); the photographic expert Snelling reported them “entirely genuine, unfaked photographs of single exposure, open-air work” with “no trace whatever of studio work involving card or paper models” (p. 53); and the lantern operator at Wakefield, “a very intelligent man who had taken a sceptical attitude, was entirely converted” by their behaviour under enlargement (pp. 91–92). They were right about the negatives, which were single exposures and unaltered, and Snelling and the operator were wrong about what was in front of the lens, since the fairies were paper cut-outs photographed in 1917. “There was little dispute that the fairies were there in front of the camera; the dispute was whether they were real or paper cutouts” (Paterson, 2013, p. 13). The examinations answered the camera’s question, whether something had been in front of the lens, and could not reach the viewer’s. In the thesis’s words, “what allowed the fairies to be viewed as real wasn’t only the quality of the render, but also the circumstances surrounding the production of the images and audience expectations of the technical capabilities of the medium” (p. 13). Belief travelled through expertise and through what the technology was believed capable of, which is to say through provenance, in 1921.
Rini argues that recordings have served as an “epistemic backstop” to testimony (Rini, 2020, p. 2), that the gravest danger of deepfakes is “not that they will trick us into believing false content, but that they will gradually eliminate the epistemic credentials of all recordings” (p. 8), and that the authority of still photographs “has already been eroded by Photoshop” (p. 13). Habgood-Coote replies that deepfakes become less concerning once “social norms’ role in recording epistemology” is recognised, and that “photographic manipulation history reveals important precedents” (Habgood-Coote, 2023). The evidence on where viewers actually take the question is Elmas’s year of r/RealOrAI, in which individuals reasoned from perceptual features 70 per cent of the time and from provenance 4 per cent, and the community’s summaries amplified provenance reasoning 4.3 times (Elmas, 2026). A viewer alone looks at the image, and viewers together ask where it came from. The institutional form of that question has moved into the camera: Leica’s M11-P is “the world’s first camera with Content Credentials built-in”, and each file records “who captured an image and when, and how they did so” (Lyon, 2023). The credential is the Kodak expert built in at capture, an examination the viewer neither performs nor could, and what it asks the viewer to trust is not the image but the maker of the camera and the process that wrote the record.
Provenance was part of the reading in 1917 and is where the question has gone now, with the 4 per cent as the measure of how far it is from answering it. While that stays unsettled, the substitution is already under way in the commercial work where the question was never asked.
What is being replaced
For blog articles and website builds I no longer license stock photographs; I write a prompt and generate the image. There is no fee and no permission to obtain, the picture fits the scenario, and the people in it are nobody, which for a stock image is the intention. None of my clients has asked whether these images were photographs, and this is substitution where photography was already only photorealism, since the stock image was never wanted as a picture of anyone in particular. Frosh describes an industry of “formulaic and stereotypical ‘generic’ images” made to “tolerate multiple reuse” (Frosh, 2003, pp. 5, 41), and a stock photograph of a woman laughing at a salad is the look of a photograph of no one, which is what was ordered. Getty’s figures are consistent with this, its creative business growing 0.7 per cent across 2025 while editorial, the coverage of events, grew 6.9 per cent (Getty Images, 2026), so that the picture of something that happened held its value better than the picture of something in general.
Fashion drew the line in public in the same year, when H&M published campaign images made with consented and watermarked “digital twins” of its models (H&M Group, 2025; Apparel Resources, 2025). A month later Guess ran models that did not exist in Vogue, disclosed in small print, and a reader said the magazine had “lost credibility” (Allison, 2025). The objection was not to generation as such but to being made to look at a person who was not one without being told. A generated image that stands in for a real person or event fails differently from a generated salad, as Amnesty found when it posted generated images of the 2021 Colombian protests and, after criticism from photojournalists, deleted them (Pontone, 2023). The rule being applied in each case was one the eye cannot apply, whether the process that made the image was the one claimed.
Where the objection falls can be predicted from what the audience valued in the image before. If an audience reads an image as the work of a person, an image that gives the same reading without the work gives the value of the reading without the value of the maker, a loss that is felt as trickery where it is felt at all. Nobody mourned the maker of a stock photograph, and nobody objected when the drone took the helicopter’s shot for the price of a lens, because the value of those images had never been in who made them. The Vogue reader objected to being shown a person who was not one, and the photojournalists objected to a record of an event that no witness had made, because in those images the maker’s presence was the value. The distinction described earlier between the constructed image and the captured one settles where the line falls. The generated image belongs on the constructed side although no hand is involved, since its accidents are the model’s, not the world’s, and not the maker’s either, as stock already was in purpose. Photography remains where the image is needed as a record, because being captured is what the record consists of.
There is nothing improper in the substitution as far as it goes, and the same tools can present a generated image as a record, which is where the old problem returns under new terms. To the question the essay opened with, whether anything but the medium has changed, the answer is that the process has not. An audience has met a new kind of image and is learning, in public and quickly, to read it, as audiences did for the buckler, the train and the render, with the change in how far the reading now has to go. The question used to be whether an image was a painting or a photograph, then whether it showed a physical space or a virtual one, and it is now whether the image was generated at all, which is asked of every image and cannot be answered from any of them. The audience that learned to see through a camera is being asked to see past one, and since looking cannot do that, it has begun, reasonably enough, to ask where the image came from instead.
References
Allison, C. (2025, July 29). Her features are flawless. But this blonde, blue-eyed model in Vogue isn’t real. ABC News. https://www.abc.net.au/news/2025-07-29/vogue-ai-model-controversy/105580372
Apparel Resources. (2025, July 3). H&M unveils first AI-generated digital model twins for ads and social media. https://apparelresources.com/business-news/manufacturing/hm-unveils-first-ai-generated-digital-model-twins-ads-social-media/
Batchen, G. (2000). Each wild idea: Writing, photography, history. MIT Press.
Bátori, Z. (2026). The creative ghost in the algorithm: AI-generated photo-based images. Synthese, 208, Article 71. https://doi.org/10.1007/s11229-026-05491-3
Bazin, A. (1960). The ontology of the photographic image (H. Gray, Trans.). Film Quarterly, 13(4), 4–9. (Original work published 1945)
Bogost, I., & Poremba, C. (2008). Can games get real? A closer look at ‘documentary’ digital games. In A. Jahn-Sudmann & R. Stockmann (Eds.), Computer games as a sociocultural phenomenon: Games without frontiers, war without tears (pp. 12–21). Palgrave Macmillan. https://doi.org/10.1057/9780230583306_2
Bottomore, S. (1999). The panicking audience? Early cinema and the ‘train effect’. Historical Journal of Film, Radio and Television, 19(2), 177–216. https://doi.org/10.1080/014396899100271
Chesney, B., & Citron, D. K. (2019). Deep fakes: A looming challenge for privacy, democracy, and national security. California Law Review, 107, 1753–1820. https://www.californialawreview.org/print/deep-fakes-a-looming-challenge-for-privacy-democracy-and-national-security
Ciment, M. (1980). Kubrick on The Shining: An interview with Michel Ciment. The Kubrick Site. http://www.visual-memory.co.uk/amk/doc/interview.ts.html (Reprinted in M. Ciment, Kubrick: The definitive edition, Faber, 2001)
DJI. (2013, October 28). DJI released Phantom 2 Vision: Your flying camera [Press release]. https://www.dji.com/media-center/announcements/dji-released-phantom-2-vision-your-flying-camera
Doyle, A. C. (1921). The coming of the fairies. George H. Doran. https://archive.org/details/comingoffairie00doyl
Eldagsen, B. (2023, April 13). Sony World Photography Awards 2023 [Statement, updated April 19, 2023]. https://www.eldagsen.com/sony-world-photography-awards-2023/
Elmas, T. (2026). Humans cannot detect AI-generated media but communities may, for now: Collaborative AI detection in r/RealOrAI on Reddit. arXiv. https://arxiv.org/abs/2605.24287
Frosh, P. (2003). The image factory: Consumer culture, photography and the visual content industry. Berg.
Geissler, D., Robertson, C., & Feuerriegel, S. (2026). Designing effective digital literacy interventions for boosting deepfake discernment. In Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems. ACM. https://doi.org/10.1145/3772318.3790428
Getty Images. (2026, March 16). Getty Images reports fourth quarter and full year 2025 results [Press release]. https://newsroom.gettyimages.com/en/getty-images/getty-images-reports-fourth-quarter-and-full-year-2025-results
Golby, J. (2023, March 27). I thought I was immune to being fooled online. Then I saw the pope in a coat. The Guardian. https://www.theguardian.com/commentisfree/2023/mar/27/pope-coat-ai-image-baby-boomers
Grohmann, L., Halle, F. A., & Appel, M. (2026). Deepfake! A liar’s dividend for audiovisual material. Psychology of Popular Media. Advance online publication. https://doi.org/10.1037/ppm0000665
Gunning, T. (1995). An aesthetic of astonishment: Early film and the (in)credulous spectator. In L. Williams (Ed.), Viewing positions: Ways of seeing film (pp. 114–133). Rutgers University Press. (Original work published 1989)
Gunning, T. (2003). Re-newing old technologies: Astonishment, second nature, and the uncanny in technology from the previous turn-of-the-century. In D. Thorburn & H. Jenkins (Eds.), Rethinking media change: The aesthetics of transition (pp. 39–60). MIT Press.
Gurevitch, L. (2010). The cinemas of interactions: Cinematics and the ‘game effect’ in the age of digital attractions. Senses of Cinema, 57. https://www.sensesofcinema.com/2010/feature-articles/the-cinemas-of-interactions-cinematics-and-the-%E2%80%98game-effect%E2%80%99-in-the-age-of-digital-attractions/
Habgood-Coote, J. (2023). Deepfakes and the epistemic apocalypse. Synthese, 201, Article 103. https://doi.org/10.1007/s11229-023-04097-3
Hagues, A. (2023, May 28). Zelda: Tears of the Kingdom: How to take pictures. Nintendo Life. https://www.nintendolife.com/guides/zelda-tears-of-the-kingdom-how-to-take-pictures
H&M Group. (2025, July 2). H&M continues its exploration of creativity with AI [Press release]. https://hmgroup.com/news/hm-continues-its-exploration-of-creativity-with-ai/
Högemann, M., Betke, J., & Thomas, O. (2025). What you see is not what you get anymore: A mixed-methods approach on human perception of AI-generated images. Frontiers in Artificial Intelligence, 8, Article 1707336. https://doi.org/10.3389/frai.2025.1707336
Ho, J., Jain, A., & Abbeel, P. (2020). Denoising diffusion probabilistic models. arXiv. https://arxiv.org/abs/2006.11239
Huang, G., & Hu, B. (2025). “A warning is not enough. Teach me how to spot deepfakes”: Testing media literacy interventions for combating deepfakes. Science Communication. Advance online publication. https://doi.org/10.1177/10755470251382889
Josipovici, G. (1998). Contre-jour: A triptych after Pierre Bonnard. Carcanet. (Original work published 1986)
Kamali, N., Nakamura, K., Chatzimparmpas, A., Hullman, J., & Groh, M. (2024). How to distinguish AI-generated images from authentic photographs. arXiv. https://arxiv.org/abs/2406.08651
Lopez, J. (2023, October 3). Celebrating one year of Stable Diffusion. Stability AI. https://stability.ai/news-updates/celebrating-one-year-of-stable-diffusion
Lyon, S. (2023, October 26). Leica launches world’s first camera with Content Credentials. Content Authenticity Initiative. https://contentauthenticity.org/blog/leica-launches-worlds-first-camera-with-content-credentials
Manovich, L. (1996). The paradoxes of digital photography. In H. von Amelunxen, S. Iglhaut, & F. Rötzer (Eds.), Photography after photography: Memory and representation in the digital age (pp. 57–65). G+B Arts.
Mehta, I. (2023, October 4). Google announces AI-powered photo editing features for new Pixel phones. TechCrunch. https://techcrunch.com/2023/10/04/google-announces-ai-powered-photo-editing-features-for-new-pixel-phones/
Midjourney [@midjourney]. (2022, July 13). We’re officially moving to open-beta! [Post]. X. https://x.com/midjourney/status/1547108864788553729
Miller, E. J., Steward, B. A., Witkower, Z., Sutherland, C. A. M., Krumhuber, E. G., & Dawel, A. (2023). AI hyperrealism: Why AI faces are perceived as more real than human ones. Psychological Science, 34(12), 1390–1403. https://doi.org/10.1177/09567976231207095
Nightingale, S. J., & Farid, H. (2022). AI-synthesized faces are indistinguishable from real faces and more trustworthy. Proceedings of the National Academy of Sciences, 119(8), Article e2120481119. https://doi.org/10.1073/pnas.2120481119
Paterson, M. (2013). Gaming and photography: Investigating the elision of illusion and the actual [Master of Design thesis, Victoria University of Wellington].
Pontone, M. (2023, May 7). Amnesty International slammed over AI protest images. Hyperallergic. https://hyperallergic.com/820339/amnesty-international-slammed-over-ai-protest-images/
Poremba, C. (2007). Point and shoot: Remediating photography in gamespace. Games and Culture, 2(1), 49–58. https://doi.org/10.1177/1555412006295397
Rini, R. (2020). Deepfakes and the epistemic backstop. Philosophers’ Imprint, 20(24), 1–16. http://www.philosophersimprint.org/020024/
Roca, T., Cintron Roman, A., Torres Vega, J., Duarte, M., Wang, P., White, K., Misra, A., & Lavista Ferres, J. (2025). How good are humans at detecting AI-generated images? Learnings from an experiment. arXiv. https://arxiv.org/abs/2507.18640
Rockstar Games. (2019b, December 13). Red Dead Redemption 2 Photo Mode and Story Mode additions now available on PS4. Rockstar Newswire. https://www.rockstargames.com/newswire/article/75o941131a8257/Red-Dead-Redemption-2-Photo-Mode-and-Story-Mode-Additions-Now-Availabl
Rombach, R., Blattmann, A., Lorenz, D., Esser, P., & Ommer, B. (2022). High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2022). https://arxiv.org/abs/2112.10752
Schiff, K. J., Schiff, D. S., & Bueno, N. S. (2025). The liar’s dividend: Can politicians claim misinformation to evade accountability? American Political Science Review, 119(1), 71–90. https://doi.org/10.1017/S0003055423001454
Schuhmann, C., Beaumont, R., Vencu, R., Gordon, C., Wightman, R., Cherti, M., Coombes, T., Katta, A., Mullis, C., Wortsman, M., Schramowski, P., Kundurthy, S., Crowson, K., Schmidt, L., Kaczmarczyk, R., & Jitsev, J. (2022). LAION-5B: An open large-scale dataset for training next generation image-text models. Advances in Neural Information Processing Systems, 35. https://arxiv.org/abs/2210.08402
Spies, T., & Urban, A. (2026). Behind the lens: Defining virtual photography through community interviews. Game Studies, 26(2). https://gamestudies.org/2602/articles/spies_urban
Steyerl, H. (2023). Mean images. New Left Review, 140/141. https://newleftreview.org/issues/ii140/articles/hito-steyerl-mean-images
Tolentino, D. (2023, March 27). AI-generated images of Pope Francis in puffer jacket fool the internet. NBC News. https://www.nbcnews.com/tech/pope-francis-ai-generated-images-fool-internet-rcna76838
Vasari, G. (1991). The lives of the artists (J. C. Bondanella & P. Bondanella, Trans.). Oxford University Press. (Original work published 1568)
Wiggers, K. (2022, July 20). OpenAI expands access to DALL-E 2, its powerful image-generating AI system. TechCrunch. https://techcrunch.com/2022/07/20/openai-expands-access-to-dall-e-2-its-powerful-image-generating-ai-system/