[Yang Myung-gyun Column] In the Era of Physical AI, Do Agricultural Data Have “Verbs”?
[Expert Column] Professor Yang Myung-gyun, Department of Bioindustrial Machinery Engineering, Jeonbuk National University
AI has moved beyond the era when it produced answers on a screen, and physical AI, which sees reality, makes judgments, and moves directly, is attracting attention. In agriculture as well, attempts to combine AI and robots for harvesting, pest control, transport, and environmental control are rapidly increasing. But before talking about robots and models, there is something we must ask first. What will we have an AI that seeks to learn behavior read?
If we were to divide agricultural data so far by part of speech, most of it would be “nouns.” Temperature, humidity, leaf area, fruit, disease symptoms, yield. But for physical AI to act in the real world, nouns alone are not enough. Agricultural data also need “verbs.”
There are many words, but few sentences
Agricultural AI has rapidly developed its ability to “see” crops. RGB Multimodal approaches that combine images with 3D, spectral·hyperspectral, thermal imaging, and environmental sensors have made it possible to express crop conditions with greater precision. Such data become the “eyes” of physical AI. That is because to act, it must first see properly.
It is not that verbs are entirely absent. Controllers in smart greenhouses record when roof vents were opened and how much nutrient solution was supplied. However, the record usually ends there. What condition the crop was in at that moment, why that judgment was made, and how it responded afterward are either elsewhere or not recorded. Tasks performed by human hands, such as leaf removal, fruit thinning, training, and harvesting, often leave no data on the actions themselves.
Ultimately, what is lacking is not only the amount of data but the connections. In what state, what was done, why it was done, and what changed as a result? Only when these four elements are connected over the flow of time do they become a single “sentence” that AI can learn from. Our agricultural data contain many words but few sentences.
Waiting is also a verb
Think of a harvesting robot. Simply collecting trajectories in which a robotic arm finds the location of fruit, approaches it, and picks it is not enough to constitute sufficient behavioral data. Why that fruit was selected, from which direction it was approached, how much force was applied, and whether there was any damage after harvest must be recorded together.
What is more interesting is the fruit that was not picked. A skilled worker sees it but does not touch it. That is because the worker has judged that it is better to wait another day or two. In agriculture, “waiting” is also an important action. But things not done usually do not remain in the record. An AI trained only on data recording what was done may learn when it should move, but it is difficult for it to learn when it should not move.
Agricultural sentences end late
In many manufacturing·logistics robot tasks, whether an object was grasped, or placed in the right spot, can be checked within a short time. In agriculture, that is often not the case. Today’s irrigation and nutrient-solution adjustments affect growth and quality days later, and leaf removal and pruning appear as differences in yield and quality weeks later. The same action can lead to different results depending on the variety, growth stage, and environment.
In sentence terms, the subject and verb are written today, but the period is placed several days or weeks later. In the meantime, the weather changes and other tasks overlap, making it difficult to determine which result came from which action. The fact that agriculture deals with living organisms is why agricultural physical AI is difficult to complete simply by importing automation from other fields as it is.
Tacit knowledge is not an answer but a hypothesis
In agriculture, people say that the experience and intuition of skilled farmers, that is, tacit knowledge, are important. The ability to look at the color and posture of leaves, and the day’s weather, and decide “let’s reduce the water a little today” certainly exists. If so, should we teach that judgment to AI as it is?
Even in the same situation, judgments may differ from one skilled farmer to another, and there is no guarantee that methods used for a long time are always the best. Good results come not only from farmers’ judgments but also from the combined effects of many conditions such as facilities, varieties, and weather. Therefore, merely collecting many records does not automatically isolate the effect of a particular judgment. Only when cases in which different actions were tried under similar conditions, or comparative experiments, are included can we verify whether that judgment was correct.
That is why tacit knowledge is closer to a good hypothesis awaiting verification than to a correct answer that AI should follow. Turning farmers’ experience into data is not a matter of copying experience as the answer, but of transforming it into knowledge that can confirm when it is valid and when it is not.
What must be asked before the name of a new model
In physical AI, VLA(vision·language·action) models, models that jointly learn from the experiences of multiple robots, and technologies that expand experience using virtual environments and world models are developing rapidly. But what must be asked first is what those models will experience and learn from.
Are we creating new agricultural data for physical AI, or are we attaching the names of new technologies to familiar data?
If until now we have recorded the state of agriculture, now it is time to record the actions of agriculture. When “verbs” such as seeing, waiting, approaching, grasping, cutting, watering, and ventilating are placed alongside nouns and connected all the way to results to form a single sentence, physical AI will finally be able to learn how to move agriculture.
This article has been automatically translated by AI (Artificial Intelligence).