Which matters are important in object recognition? - Chapter 6
- What are the computational problems in object recognition?
- What are the multiple pathways for visual perception?
- What are the representational differences between the dorsal and ventral streams?
- What is the difference between perception for identification and perception for action?
- How do you see shapes and perceive objects?
- What is the role of the grandmother cell in ensemble coding?
- What are the top-down effects of object recognition?
- Can you read minds?
- How does the specificity of object recognition in higher visual areas work?
- What are failures in object recognition?
- What is prosopagnosia?
What are the computational problems in object recognition?
When you think about object recognition, there are a few things to keep in mind:
Use terms precisely: When you talk about certain cases or patients it is very important for researchers to be precise about when using terms like perceive or recognize.
Object perception is unified: our sensory system uses a divide-and-conquer strategy, but the perception of objects is unified. Features like color and motion are processed along distinct neural pathways. Perception, however, requires more than simply perceiving the features of objects.
Perceptual capabilities are enormously flexible and robust: The city vista looks the same whether the view of both eyes or with only the left or the right eye. The percept of an image always stays the same, even if we stand on our head and the retinal image is inverted.
The product of perception is immediately interwoven with memory: object recognition is more than linking features to form a coherent whole. Part of memory retrieval is recognizing that things belong to certain categories.
Object constancy refers to our amazing ability to recognize an object in countless situations. When you show a drawing of a car, from a different view each time, a person has no problem identifying the object in each picture as a car, and discerning that all four cars are the same model. The visual information emanating from an object varies as a function of three factors: viewing position, illumination conditions, and context.
Viewing position: sensory information depends highly on your viewpoint, which changes not only as you view an object from different angles, but also when the object itself moves and thus changes its orientation relative to you. The human perceptual system is adept at separating changes caused by shifts in viewpoint from changes intrinsic to an object itself. The sensory system automatically uses any sensory cues and past knowledge to maintain object constancy.
Illumination: while the visible parts of an object may differ depending on how light hits it and where shadows are cast, recognition is largely insensitive to changes in illumination. A dog in the sun and a dog in the shade both register as a dog.
Context: objects are rarely seen in isolation. People see objects surrounded by other objects and against varied backgrounds. Yet we have no trouble separating, for instance, a dog from other objects on a crowded city street. Our perceptual system quickly partitions the scene into components.
Object recognition must accommodate these three sources of variability. But the system also has to recognize that changes in perceived shape may reflect actual changes in the object.
What are the multiple pathways for visual perception?
The pathways carrying visual information from the retina to the first few synapses in the cortex segregate into multiple processing streams. Much of the information goes to the V1 (primary visual cortex). Output from the V1 is contained primarily in two major fiber bundles, which carry visual information to regions of the parietal and temporal cortex - that are involved in visual object recognition.
There is the ventral (occipitotemporal) stream and the dorsal (occipitoparietal) stream. These two are also known as the what and where pathways. Ungerleider and Mishkin proposed hat processing along these two pathways is designed to extract fundamentally different types of information.
They hypothesized that the ventral stream is specialized for object perception and recognition, for determining 'what' we are looking at. The dorsal stream is specialized for spatial perception, for determining 'where' an object is, and for analyzing the spatial configuration between different objects in a scene. What and where are two basic questions to be answered in visual perception. To be more specific, the dorsal stream or 'where' is going upwards from the back of the brains to the front, the ventral stream is 'what' and is going downwards from the back of the brain to the front. If you want to use a mnemonic; dorsal comes first in the alphabet when you choose between dorsal and ventral, so dorsal is above and ventral is below.
The first data for the what-where dissociation of the ventral and dorsal stream comes from animal studies. Animals with bilateral lesions to the temporal lobe that disrupted the ventral stream had great difficulty discriminating between different shapes (the 'what' discrimination). But, these animals had no problems determining where the object was in relation to other objects, because this second ability depends greatly on the 'where' route. But the separation of the 'what' and 'where' routes is not limited to the visual system, for instance in audition.
What are the representational differences between the dorsal and ventral streams?
Neurons in both the temporal and parietal lobes have large receptive fields, but the physiological properties of the neurons within each lobe are quite distinct. 40% of these neurons have receptive fields near the central region of vision, the remaining cells have receptive fields that exclude the foveal region. These eccentrically tuned cells are ideally suited for detecting the presence and location of a stimulus, especially one that has just entered the field of view.
The response for neurons in the ventral stream of the temporal lobe is quite different. The receptive fields for these neurons always encompass the fovea, most of these neurons can be activated by a stimulus that falls within either the left or the right visual field. Cells within the visual areas of the temporal lobe have a diverse pattern of selectivity. In the posterior region cells show a preference for relatively simple features such as edges. Further along the process stream, they have a preference for much more complex figures; such as human body parts, apples, flowers or snakes etc.
What is the difference between perception for identification and perception for action?
Agnosia is an inability in processing sensory information even though the sense organs and memory are not defective. To be agnosic means to experience a failure of knowledge or recognition of objects, persons, shapes, sounds, or smells. When the disorder is limited to the visual modality, is it referred to as visual agnosia. This is a deficit in recognizing objects even when the processes for analyzing basic properties such as shape, color, and motion are relatively intact.
Patient DF is an extraordinary case. She couldn't name the right household items, made errors in labeling them. She usually gave crude descriptions of displayed objects. Picture recognition was even more disrupted. When DF was given an explicit matching task she failed miserably. She couldn't orientate the card the right way for fitting the lock. But when she was asked to insert the card into the slot, DF quickly reached forward and inserted the card into the lock. The explicit matching task couldn't succeed because DF could not recognize the orientation of the object because of the severe agnosia. But when DF was asked to insert the card, the shape and orientation information were available for the visuomotor task. So, the 'where' system appears to be essential for more than determining the locations of different objects; it is also critical for guiding interaction with these objects.
So, patients with selective lesions in the ventral pathway may have severe problems in consciously identifying objects, yet they can use the visual information to guide coordinated movement. Thus we see that visual information is used for a variety of purposes.
Optic ataxia holds that patients can recognize objects, yet they cannot use visual information to guide their actions. When someone with optic atraxia reaches for an object, she doesn't move directly toward it; rather, she gropes about like a person trying to find something in the dark. Optic atraxia is associated with lesions in the parietal cortex.
How do you see shapes and perceive objects?
Object perception depends primarily on an analysis of the shape of a visual stimulus, though cues such as color, texture and motion certainly also contribute to normal perception. But, even when the surface features are absent or applied inappropriately (think about an abstract painting), we are still able to recognize the object using perceptual ability to match the analysis of shape and form to an object, regardless of color, texture, or motion cues.
One way to investigate how we encode shapes is to identify areas of the brain that are active when we compare contours that form a recognizable shape versus contours that are just squiggles. There is an idea that perception involves a connection between sensation and memory in the brain. Researchers explored this question using a PET study designed to isolate the specific mental operations used when people viewed familiar shapes, novel shapes, or stimuli formed by scrambling the shapes to generate random drawings. Viewing both novel and familiar stimuli led to increases in regional cerebral blood flow bilaterally in lateral occipital cortex (LOC). Many others have also shown that the LOC is critical for shape and object recognition. People have an insensitivity to the specific visual cues that define an object, this is known as cue invariance. Thus, the LOC can support the perception of an elephant even when the elephant is blue and green, or an apple shape even when the apple is made of onyx and striped.
The functional specification of the LOC can also be tested with 6-month-old babies. To do this researchers use a fNIRS, functional near-infrared spectroscopy, which employs a lightweight system that looks similar to an EEG cap and can be comfortably placed on the infant's head. This system uses infrared light, that can project through the head and skull. The repetition suppression (RS) effect is hypothesized to indicate increased neural efficiency: the neural response to the stimulus is more efficient and perhaps faster when the pattern has been recently activated.
From shapes to objects
Multistable perception is an image where their is an object that you can see in a black-and-white view, such as a vase, but when you point your attention to another part of the image you see a different object. The vase can change profiles in two people facing each other, an then you can go back to the vase, back to the two people, on and on. This is an example of multistable perception. The stimulus information does not change at the points of transition from one percept to the other, but the interpretation of the pictorial cues does.
What is the role of the grandmother cell in ensemble coding?
How do we recognize specific objects? Are there individual cells that respond only to specific integrated percepts, or does perception of an object depend on the firing of a collection or ensemble of cells? In the latter case, this would mean that when you see a peach, a group of neurons that code or different features of the peach might become active, with some subset of them also active when you see a nectarine. A type of neuron that can recognize a complex object is called a gnotic unit, referring to the idea that the cell signals the presence of a known stimulus - an object, a place, or an animal that has been encountered in the past.
Researchers also discovered cells in the IT gyrus and the floor of the superior temporal sulcus (STS) that are selectively activated by faces. They coined the term grandmother cell to convey the notion that people's brains might have a gnostic unit that becomes excited only when their grandmother comes into view. Although it is tempting to conclude that there are cells like this that are gnostic units, it is important to keep in mind the limitations of such experiments:
Aside from the infinite number of possible stimuli, the recordings are performed on only a small subset of neurons. This cell potentially could be activated by a broader set of stimuli, and many other neurons might respond in a similar manner.
The results also suggest that these gnostic-like units are not really 'perceptual'. The cell could represent a concept of, for instance, a 'grandmother'.
One alternative to the grandmother-cell hypothesis is that object recognition results from activation across complex feature detectors. Granny, then, is recognized when some of these higher-order neurons are activated. According to this ensemble hypothesis, recognition is not due to one unit but to the collective activation of many units.
What are the top-down effects of object recognition?
Up to this point, we have emphasized a bottom-up perspective on processing within the visual system, showing how a multilayered system can combine features into more complex representations. This model appears to nicely capture the flow of information along the ventral pathway. But there is also an top-down way we are not forgetting. One model of top-down effects emphasizes that input from the frontal cortex can influence processing along the ventral pathway. The frontal lobe generates predictions about what the scene is, using this early scene analysis and knowledge of the current context. These top-down predictions can then be compared with the bottom-up analysis occurring along the ventral pathway of the temporal cortex, making for faster object recognition by limiting the field of possibilities.
Can you read minds?
We have seen various ways in which scientists have show us that you can manipulate the output and input of the visual cortex. These observations have led investigators to realize that it should, at least in principle, be possible to analyze the system in the opposite direction. That is, we should be able to look at someone's brain activity and infer what the person is currently seeing - a form of mind reading. This idea is referred to as decoding: the brain activity provides the coded message, and the challenge is to decipher it and infer what is being represented.
There are two issues:
Our ability to decode mental states is limited by our models of how the brain encodes information.
Our ability to decode will be limited by the resolution of our measurement systems.
How does the specificity of object recognition in higher visual areas work?
When we meet someone, we always look at that person's face. The face, particularly the eyes, of another person can provide significant cues about what is important in his environment. Also, looking at someone's lip when they are speaking can provide a lot more information about what that person is saying.
Is face processing special?
It seems reasonable to suppose that our brains have a general-purpose system for recognizing all sorts of visual inputs, with faces constituting just one important class of problems to solve. But multiple studies argue that face perception does not use the same processing mechanisms as those used in object recognition, but instead depends on a specialized network of brain regions. Do the processes of face recognition and nonfacial object recognition involve physically distinct mechanisms? Although clinical evidence showed that people could have what appeared to be selective problems in face perception, more compelling evidence of specialized face perception mechanims comes from neurophysiological studies with nonhuman primates. Neurons in various areas of the monkey brain show selectivity for face stimuli.
The similar specificity for faces is observed using fMRI studies in humans, including an area in the right fusiform gyrus parahippocampal place area (PPA). This area is specialized for processing information about spatial properties, for instance the difference between an indoor and outdoor scene, and the extrastriate body area (EBA) and the fusiform body area (FBA) have been identified as more active when body parts are viewed.
What are failures in object recognition?
Patients with visual agnosia have provided a window into the processes that underlie object recognition. By analyzing the subtypes of visual agnosia and their associated deficits, we can draw inferences about the processes that lead to object recognition. Although the term visual agnosia has been applied to a number of distinct disorders associated with different neural deficits, patients with visual agnosia generally have difficulty recognizing objects that are presented visually or require the use of visually based representations.
The current literature broadly distinguishes between three major subtypes of visual agnosia: apperceptive, integrative and associative.
Apperceptive visual agnosia: The recognition problem is one of developing a coherent percept: the basic components are there, but they can't be assembled. It's somewhat like going to Legoland, but instead of seeing buildings, cars, and monsters, you can only see piles of Lego bricks. The elementary visual functions - acuity, colorvision, and brightness - are still intact. The object recognition problems become especially evident when a patient is asked to identify objects on the basis of limited stimulus information, for instance when the object is shown as a line drawing or is seen from an unusual perspective.
Integrative visual agnosia: this is a subtype of the apperceptive visual agnosia, where people perceive the parts of an object but are unable to integrate them into a coherent whole. At Legoland they may see walls and windows, but not a house. A patient's object recognition problems became apparent when he was asked to identify objects that overlapped each other.
Associative visual agnosia: perception occurs without recognition. It is the inability to link a percept with its semantic information, such as its name, properties or functions. A patient can perceive objects with het visual system but cannot understand them or assign meaning to them. At Legoland she may perceive a house, and be able to draw a picture of that house, but still be unable to tell that it is a house or describe what a house is for.
Patients with agnosia are unable to recognize common objects. This deficit is modality specific. Patients with visual agnosia can recognize an object when they touch, smell, taste or hear it, but not when they can only see it. Therefor, visual agnosia can be category specific. Category-specific deficits are deficits of object recognition that are restricted to certain classes of objects. Linked to this there has been a debate in research about how object knowledge is organized in the brain. One theory suggests that it is organized by features and motor properties, and the other suggests specific domains relevant to survival and reproduction.
What is prosopagnosia?
Prosopagnosia is the term used to describe an impairment in face recognition. Given the importance of face recognition, propsopagnosia is one of the most fascinating and disturbing disorder of object recognition. Propsopagnosia is usually observed in patients who have lesions in the ventral pathway, especially occiptial regions associated with face perception and the fusiform face area. Some patients also have congenital propsopagnosia (CP), defined as a lifetime impairment in face recognition that cannot be attributed to a known neurological condition.
Hostilic processing is a form of perceptual analysis that emphasizes the overall shape of an object. This mode of processing is especially important for face perception. We can recognize a face by the overall configuration of its features, and not by the individual features itself.
Analysis-by-parts processing is a form of perceptual analysis that emphasizes the component parts of an object. This mode of processing is important for reading, when we decompose the overall shape into its constituent parts.
Join with a free account for more service, or become a member for full access to exclusives and extra support of WorldSupporter >>
Concept of JoHo WorldSupporter
JoHo WorldSupporter mission and vision:
- JoHo wants to enable people and organizations to develop and work better together, and thereby contribute to a tolerant and sustainable world. Through physical and online platforms, it supports personal development and promote international cooperation is encouraged.
JoHo concept:
- As a JoHo donor, member or insured, you provide support to the JoHo objectives. JoHo then supports you with tools, coaching and benefits in the areas of personal development and international activities.
- JoHo's core services include: study support, competence development, coaching and insurance mediation when departure abroad.
Join JoHo WorldSupporter!
for a modest and sustainable investment in yourself, and a valued contribution to what JoHo stands for
Work for JoHo WorldSupporter?
Volunteering: WorldSupporter moderators and Summary Supporters
Volunteering: Share your summaries or study notes
Student jobs: Part-time work as study assistant in Leiden
- Insurance for emigrants, expats and living abroad: international insurance for expats and emigrants
- Insurance for activities abroad: Backpacking Travel abroad Intern abroad Study abroad Volunteer abroad Work abroad
- Insurance: ACS Globe Traveller Caremed Insurances Expatriate Travel Insurance IMG’s GlobeHopper World Nomads Insurance SafetyWing Insurance JoHo Special ISIS verzekering NL/BE Working Nomad verzekering NL/BE More about Insurance for abroad
Search only via club, country, goal, study, topic or sector
Select any filter and click on Search to see results








