Type classification for lemmas
Every lemma carries exactly one type. The type vocabulary has two sorts: ontological types, for lemmas that encode a concept, and procedural types, for lemmas that encode an instruction for processing. Which sort a lemma takes is the first decision the annotator makes; everything else follows from it.
A lemma here is a lexeme: a citation form paired with a part of speech, not a bare string. A lexical unit (LU) is a lemma paired with a frame, carrying a sense description — the relation is N:M, exactly as in Berkeley FrameNet. The type belongs to the lemma and only to the lemma. The LU inherits it and carries no type of its own, so creating an LU involves no typing decision at all.
What this layer claims, and what it does not#
Four scoping commitments govern every assignment in this document.
- The type is a property of the lemma — a lexeme, text plus POS — not of a word form or a referent. POS already does the construal-fixing work: colar/VERB and colar/NOUN are different lexemes with different meanings before any frame is consulted.
- The type classifies a conceptual profile, not the furniture of the world. No claim is made about what exists. This follows the grounding ontology: DOLCE is a Descriptive Ontology, an inventory of particulars as conceptualized by natural language, explicitly neutral on metaphysics. Typing a lemma against DUL types a conceptualization against a catalogue of conceptualizations.
- The type is a coarse, defeasible default, serving as the input to constructional composition and coercion. A construction may override it — override is the interesting data, not a counterexample, since a function needs an argument. The default lives in the lexicon and is defeated in use.
- The type is a comparative concept, in Haspelmath's sense: an FNBr-internal instrument for cross-frame, cross-resource and cross-linguistic alignment — not a descriptive category of Portuguese, and not a metaphysical claim.
Because the type is defeasible (commitment 3), the class-compatibility check in §8 flags a marked combination worth review, not an error.
1. Parts of speech and the locus of the type#
1.1 Why retain POS at all#
The theoretical case against parts of speech as a semantic category is strong: syntactic categories are language-particular and construction-derived (Croft, Radical Construction Grammar), and do not translate across languages as descriptive categories (Haspelmath 2010). This is why the type layer exists and why it is not a POS layer.
POS is nonetheless retained: it is load-bearing for lemmatization and morphological analysis, MWE and terminology identification, corpus query (word sketches are defined over POS patterns), and interoperability with UD, LMF and OntoLex-Lemon/LexInfo.
The architecture follows OntoLex-Lemon: POS is morphosyntactic metadata on the lexical entry; the type is the conceptual layer, linked by semantics-by-reference. Both sit on the lemma; neither pretends to be the other.
1.2 POS is part of lemma identity#
POS is a property of the lexical entry = (lemma string + POS), and the lemma key must include it:
lemma(id, text, pos, lang, homograph_idx, type_id)
UNIQUE(text, pos, lang, homograph_idx)
form (id, text, lemma_id, feats) -- the paradigm of ONE lemma entry
lu (id, lemma_id, frame_id, sense_description)
Keying lemma on text alone conflates distinct lexemes. Colar is two
lemmas: colar/VERB ("to glue"), whose paradigm contains colo, colei,
colando; and colar/NOUN ("necklace"), whose paradigm is colar,
colares. With a text-only key, a parse of Eu colo o papel would resolve
colo to a single colar row and make the noun sense reachable from a verb
form — a lemma that does not exist.
Two consequences:
- Form lookup is set-valued.
form.text → {lemma_id}returns candidates. POS is what filters them, but POS lives on the token, not on the form string: parsing the sentence assigns the token a UPOS, and matching that against the POS carried by each candidate lemma picks the right one. Colo is itself ambiguous off any list — it is a form ofcolar/VERBand also a noun in its own right (colo/NOUN, "lap") — so the string alone does not resolve it; the token's POS from the parse does, no semantic reasoning needed. - When UPOS is absent or unreliable, do not guess at the lemma layer. Return the candidate set and let frame matching decide. POS narrows; it does not adjudicate.
homograph_idx distinguishes lemmas that share text and POS: genuine
homonymy (manga fruit/sleeve) and type divergence, where one form
under one POS carries readings of different types. Default 0. It is not
used for polysemy within a single type — that is what LUs are for (§1.4).
For MWEs, pos is X and entry_type ∈ {word, mwe, affix} carries the
word/MWE/affix distinction. UD's own X tag is defined for "words that for
some reason cannot be assigned a real part-of-speech
category" — exactly the case
for dar um jeito. It asserts that no single POS applies, rather than
asserting a wrong one. This is also where the type layer outperforms POS: an
MWE has no coherent part of speech but a perfectly coherent ontological type.
X, not null. A null POS propagates: view_lu derives an LU's display
name as concat(lemma.name, udpos.suffix), and concat() returns null when
any argument is — so a null POS silently blanks the rendered name. X carries
a real suffix, so dar um jeito renders as dar um jeito.x. Every lemma
carries a POS; a lemma arriving without one is assigned X rather than left
null. (lemma.idUDPOS remains nullable in the schema for legacy reasons, but
nothing relies on that.)
1.3 Why the type belongs to the lemma#
Three reasons, one theoretical and two practical.
The lexeme is already a form–meaning pairing. Once POS is part of lemma
identity, most of the construal-fixing work is already done before any frame
is consulted. Colarᵥ and colarₙ are separate lemmas with separate types;
so are dependenteₙ (role) and dependenteₐ (value). Putting the type on
the lemma–frame pairing instead would record, on every LU, a fact constant
across all of a lemma's LUs.
The type is available before frame assignment. A parser can type a token as soon as it resolves the lemma, with no frame in hand — which is precisely when a type is most useful as a filter.
It matches the FrameNet data model. The LU table is a lemma × frame association with a sense description. A mandatory type column there would make the LU a second lexical entry rather than a pairing.
1.4 One lemma per (text, POS, type)#
Where one form under one POS carries readings of different types, each
type gets its own lemma row, distinguished by homograph_idx. Readings that
share a type stay in one lemma and are distinguished as LUs.
One lemma per (text, POS, type).
This is the only mechanism for type divergence in the specification — there is no LU-level type and no override.
- Creating an LU involves no typing decision. The annotator pairs a lemma with a frame and writes a sense description; the type arrived with the lemma.
lemma.type_idis correct by construction. It is non-null, single-valued and never wrong for any LU beneath it, because a divergent reading is by definition a different lemma.
The split is bounded by type, not by sense count: de has at least six distinguishable readings but only two types, so it yields two lemmas, not six (§6.2).
Three case shapes are known:
- Cross-sort polysemy. Negação names the act of denying (
event, in an@agentiveframe) and the propositional operator (polarity) →negação₀event,negação₁polarity. - Product nominalization. Pintura as the activity is
event; as the canvas on the wall it isobject. Same for construção, tradução, invenção, criação, publicação — deverbal nominals whose event produces an artifact that outlives it. Not all deverbal nominals: see §4. - Adpositions. De, a, em, por, com, para, sobre — the largest group, worked through in §6.2.
Not for:
- Construal shifts. A lemma can be referred to, or otherwise used, in a
way that leaves part of its normal profile inactive without the stored
type changing — the type is a fact about the lemma; the construal is a
fact about the occurrence. Pai used to pick out an individual with no
relation in play is still the
rolelemma: the relation is simply not invoked by that use. Löbner's determiner-triggered readings are the best-studied case, but the mechanism is not limited to determiners — see the next bullet. Per commitment 3, a construal shift is defeasance in use, not a second type — and, since nothing about the lemma itself changed, not a second lemma either. - Substantivized adjectives used elliptically. O bonito, um
americano pick out an entity the adjective characterizes — a person, a
car, whatever the covert head recovers as — via the construction that lets
any adjective head an NP under a determiner with its head noun elided.
This is a construal shift like the one above: bonito and americano
stay
value, and which entity the value is predicated of is a frame-level fact (§7), not a lemma-level one. The test: can the use be regenerated as "[some contextually recoverable sortal] that is ADJ," with the adjective keeping its ordinary value? If the covert head is freely recoverable and swappable by context, it is a construal shift and the adjective lemma is untouched. If the "noun" reading is fixed regardless of context and carries content the adjective does not — typically a specific institutional or relational profile — the use has lexicalized, and §1.2's POS split applies instead: dependenteₙ ("meu dependente", arole— someone for whom a specific provider must exist) is a different lemma from dependenteₐ ("ele é dependente",value), not a construal shift of it (§7). - Polysemy within a type. Genitive, material and partitive de are one
predicationlemma with several LUs. Two frames do not imply two lemmas; two types do.
Splitting is the same operation as homonymy handling, deliberately:
manga fruit/sleeve and de₀/de₁ produce identical rows. Annotators need
not decide whether two readings are "the same word" — once the type is
settled, that judgement has no consequence for the data.
A split records no relation between the two lemmas. Where relatedness
matters — negação act/operator, pintura activity/canvas — it belongs at
the lexical layer as an explicit link (OntoLex-Lemon's vartrans module),
not as a shared row.
1.5 POS and type are not orthogonal#
Deverbal nominalizations and participles show strong type/POS correlation, because category is assigned in a structure that also determines event structure. The claim is narrower: the type is not derivable from POS and not reducible to it. How much independent information it carries is answerable by a POS × type contingency table, since both columns live on the same row.
2. The two sorts: conceptual and procedural encoding#
Not every word encodes a concept. Quem, não, portanto, o have meanings, but those meanings are not things classifiable as objects, events or qualities — the "world thing" is computed in context, if at all.
Relevance Theory supplies the distinction. Blakemore (Semantic Constraints on Relevance, 1987) and Wilson & Sperber ("Linguistic form and relevance", Lingua 1993) distinguish conceptual encoding — the expression contributes a constituent to the proposition expressed — from procedural encoding — the expression encodes an instruction for processing: a constraint on inference, reference resolution, or how to build or access a representation. Discourse connectives, indexicals, polarity, and most grammatical morphemes are procedural on this analysis.
FNBr adopts the distinction as the top branch of the type vocabulary:
| Sort | Encodes | Gets an ontological type | Gets a DUL reference |
|---|---|---|---|
| conceptual | a concept | yes (§3) | yes |
| procedural | an instruction for processing | no | no |
Procedural lemmas carry no DUL reference — there is nothing in the conceptualization inventory for them to reference, because they name an operation on a conceptualized particular, not the particular itself.
3. The ontological types#
Six types. Each grounds in a DUL class. There are no parameters.
| Type | DUL grounding | Criterion | Examples |
|---|---|---|---|
event |
dul:Event |
occurs or happens: it exists in time, has (or can have) temporal parts, and does not persist unchanged | destruição, aquecer, correr, quebrar, derreter, existir |
object |
dul:Object |
a sortal — it supplies a criterion of identity — whose instances need no other entity to exist | cadeira, água, pessoa, documento, lei, ideia, criança, adulto |
role |
dul:Role ⊑ dul:Concept |
a sortal whose instances are externally dependent: some other entity, not a part of them, must exist | comprador, jogador, cliente, filho, pai, proprietário |
relation |
dul:Relation ⊑ dul:Description |
a reified relation type — a Description that defines the argument positions (roles) it relates | parentesco, paternidade, posse, casamento, adjacência, anterioridade |
quality |
dul:Quality |
a dimension — a gradable or categorial axis — that inheres in some bearer | cor, tamanho, beleza, inteligência, temperatura, cansaço, postura |
value |
dul:Region ⊑ dul:Abstract |
a region or point within a quality's dimension, nameable without presupposing any bearer | azul, quente, bonito, inteligente, cansado, quebrado, sentado, belamente |
The criteria draw on rigidity, sortality, external dependence, and the dimension/region distinction — defined in full in Appendix §A.1.
The reference principle#
The lemma references a conceptualization, never a particular and never a formal object:
- comprador references
dul:Role— a Concept that classifies persons — not any person. - parentesco references
dul:Relation, a Description — the relation type — not any particular kinship instance. - adjacência and anterioridade reference the commonsense
dul:Relation, not the formal counterpart underdul:FormalEntity.
This is what "semantics by reference" means for a lexicon: the lemma points at a conceptualization, which is exactly what a word does.
role and relation are two halves of one structure#
dul:Relation is a Description; the Concepts filling its argument positions
are dul:isDefinedIn that Description:
| FNBr object | DUL | DnS relation |
|---|---|---|
| relational frame | dul:Relation ⊑ Description |
— |
| FE of that frame | dul:Role / dul:Concept |
dul:isDefinedIn → the Description |
relation lemma (parentesco) |
references the Relation |
— |
role lemma (filho, pai) |
references a Role |
isDefinedIn the same Relation |
Parentesco names the description; filho and pai name Concepts it defines — not two unrelated primitives that happen to sound alike.
The instance layer, where needed, is dul:Situation: the particular
fatherhood holding between two individuals in an annotated sentence is a
Situation that dul:satisfies the Relation. Lemma layer → Description;
annotation layer → Situation — the same coarse-identity/fine-situation
asymmetry the whole model runs on (§7).
Converse pairs fall out of the structure. Both role and relation are
grounded in DnS, so both pai and filho are isDefinedIn the same
parentesco Relation, and the converse pairing is readable off that link.
Rigidity is not used as a typing axis, and abstract nouns do not get a type of their own — both are informational notes rather than operative rules; see Appendix §A.2 and §A.3.
4. Conceptual diagnostics#
The first question: conceptual or procedural?#
Does the lemma contribute a constituent to the proposition, or instruct the hearer how to process one? If the latter, go to §5 and stop.
object / role#
Both are thing-like; one binary test separates them: external dependence.
| Type | Examples | |
|---|---|---|
| not externally dependent | object |
pessoa, cadeira, documento, criança, adulto |
| externally dependent | role |
comprador, filho, presidente, proprietário |
Is there an entity y, not a part or constituent of x, whose existence is required for x to be P?
Comprador → yes (a purchase, a seller). Filho → yes (a parent).
Documento → no. Criança → no: being a child requires no second entity, so
it is object, not role, even though an individual stops being one
(Appendix §A.2).
Mere participation is not dependence. Documents get signed, filed and
sent, but no particular event must exist for a document to be a document. If
participation sufficed, every entity noun in the lexicon would be a role.
Two corpus diagnostics, for cases where the ontological test is hard to apply — both operationalize Löbner's [R] (Appendix §A.1):
- Bare predicative:
?Ela é irmãvsEla é médica. A +R concept resists the bare predicative without a recoverable relatum. - De-genitive reading: does de X give an argument (pai de João) or a modifier (carro de João)? Argument readings indicate +R.
These are annotatable and scorable, so the role boundary can be validated
rather than defended by introspection.
event / object in deverbal nominals#
A deverbal noun is an event by default and stays one. Nominalization is a
constructional fact: it lets an event concept be referred to, counted,
determined and modified like a thing, without changing what the concept is.
Duas destruições, a destruição foi total, aquela destruição are all the
event under nominal packaging. Countability, determination, pluralization
and adjectival modification are never evidence for object.
The one case that diverges is a lemma with a product sense: a participant of the event — typically its Result or Product — with existence independent of the event, persisting after it ends.
Does the lemma have a sense denoting a thing that the event produced, rather than the event itself under some construal?
Three tests, all passed by the product sense and failed by the event sense:
- Spatial location. A pintura está na parede. / *A destruição está na parede.
- Physical or material predicates. A tradução tem 300 páginas. / *A destruição tem 300 páginas.
- Persistence. The product exists once the event is over; the event does not.
Pintura, construção, tradução, invenção, criação, publicação pass. Destruição, aquecimento, chegada, corrida, viagem fail: they have no product, and their apparent "result" readings are the event's Result or Effect FE reached metonymically — a frame-layer fact, not a type-layer one.
Grimshaw's complex-event/simple-event/result trichotomy is not used here: it separates nominals by argument structure, not by type (Appendix §A.4).
quality / value#
Ask whether the lemma names the dimension or a region within it (Appendix §A.1). Altura/alto, temperatura/quente, beleza/bonito, cansaço/cansado are each a dimension/value pair, decided the same way every time.
No entity is presupposed by either. A temperature scale exists, and a
point on it is nameable, whether or not any object is currently at that
temperature — está quente hoje, said of the weather, names no entity at
all. What links a value to the entity bearing it is the job of the
Attributes schema (its Entity, Value and Attribute elements), not of
the lemma's type.
The type does not encode causal origin. Whether a quality resulted from
an event (cansaço, force-maintained postura) or not (beleza,
inteligência) plays no role in the classification: both are quality, and
their values — cansado, quebrado, sentado on one side, bonito,
inteligente on the other — are both value. That distinction belongs to
the situation, which the frame layer already records (a STATIVE schema
vs an Attributes schema plus a relation to the producing event); it is the
same asymmetry already accepted for destruição, which stays event while
its frames range over @agentive, @change and @stative without the
lemma splitting. See Appendix §A.5 for the full argument, including why no
reliable test separates a "resultant" quality from a non-resultant one.
Referential use does not create a new type. O bonito, um americano refer to an entity characterized by the value without the lemma becoming a different kind of thing — see §1.4, "Substantivized adjectives used elliptically," for the construal-shift diagnostic that tells this apart from genuine lexicalization (dependenteₙ/dependenteₐ, §7).
relation / role#
The tie is relation (paternidade, posse); the argument position
named as such is role (pai, proprietário). Arity and converse
evidence (pai-de ↔ filho-de, posse ↔ pertence-a) confirm a relation
exists; the link between the two is dul:isDefinedIn (§3), not a parameter.
A free-standing referential use of a role noun — pai picking out the
individual who walked in, with no relation in play — is a construal shift,
not a second type. The lemma remains role (§1.4).
5. Procedural types#
Nine types. Procedural lemmas take no DUL reference and are exempt from the class-compatibility check (§8).
| Type | Encodes | Examples |
|---|---|---|
reference |
how to resolve a referent — deictic, anaphoric, definiteness | o, este, ele, quem, aqui, hoje, agora |
quantification |
how to quantify over a domain | todo, cada, algum, vários, ambos, sempre |
polarity |
assert or deny instantiation | sim, não, nunca, jamais, nada, nenhum, sem |
degree |
how to locate a value relative to a standard | muito, pouco, mais, menos, bastante, demais, tão |
modality |
epistemic or deontic status of the proposition, with no scale named | talvez; modal auxiliaries (poder, dever, ter que); free mood markers |
connection |
inferential or discourse relation between units | contudo, portanto, porém, entretanto, pois, se |
focus |
information-structural instruction | só, também, até, mesmo, inclusive |
predication |
how to build the predication; contributes no content | copulas (ser, estar), auxiliaries, support verbs, case-marking de/a/por |
interaction |
how to manage the exchange — appeal, stance display, information source, hedging | né, hein, viu, tá, ué, nossa, olha, tipo, assim, diz que |
The nine types are organized by functional category; their theoretical provenance and relation to Relevance Theory is in Appendix §A.6.
Notes on the boundaries#
degree vs value. The criterion is whether the lemma names a region
or instructs where to locate one relative to a standard. Quente,
morno, azul name regions → value. Muito, pouco, bastante locate a
region against a contextually supplied standard → degree.
Prepositions are not a single case. ADP splits across both sorts:
contentful spatial and temporal adpositions are relation, governed
case-markers are predication, and some are polarity or connection
(§6.2).
modality vs value: does the lemma name a scale? Epistemic adverbs
with a lexicalized dimension are value, not modality. Provavelmente,
certamente, possivelmente and seguramente each name a region on the
probability/certainty dimension (probabilidade, certeza), have a gradable
adjectival counterpart (provável, certo), and stand to it exactly as
belamente stands to beleza (§6.1). Three tests separate them from
modality:
| Test | value — provavelmente |
modality — talvez |
|---|---|---|
degree modification |
muito provavelmente ✓ | *muito talvez ✗ |
| lexicalized dimension and adjective | probabilidade, provável ✓ | none ✗ |
| embeds under negation | não é provável que… ✓ | *não talvez… ✗ |
| governs mood | no | yes — subjunctive |
The first test is decisive on its own: degree modifies value by
definition, so anything muito can stack on is a value. Talvez fails all
three and governs the subjunctive — an instruction to the grammar rather than
a region on a scale. Modal auxiliaries form a third group with no scale at
all and stay in modality.
The copula and support verbs are predication. They instruct how to
assemble a predication and contribute no concept of their own.
interaction is about the exchange, not the proposition. It covers
interjections (ué, nossa, ai), interactional and tag particles (né,
hein, viu, tá), attention markers (olha, escuta), hedges and
approximators (tipo, assim), and quotative or evidential markers (diz
que, dizem que) — whatever manages the interaction rather than
contributing to or operating on the proposition expressed.
Three boundaries:
- vs
connection. A connective relates two propositions inferentially (portanto, contudo); an interactional particle relates the utterance to the hearer (né seeks confirmation, viu secures uptake). Test: does removing it leave an inferential gap between two units, or only a blunter utterance? - vs
modality.modalitymarks the speaker's own commitment to the proposition (talvez, poder); evidentials mark where the information came from (diz que). FNBr keeps evidentials ininteractionbecause they invoke the hearer's assessment of the source rather than stating the speaker's confidence (Appendix §A.7). - vs
focus. Só, até, mesmo operate on the proposition's information structure and arefocus. Tipo and assim as hedges signal approximation in the exchange rather than restructuring the proposition.
interaction lemmas evoke pragmatic frames like any other lemma; they are
exempt from the class-compatibility check of §8 along with the other
procedural types.
6. Parts of speech the type layer cuts across#
Some UPOS tags collect lemmas with almost nothing semantically in common. The
type layer is expected to cut across them; ADV and ADP need worked
guidance.
6.1 Adverbs#
"Adverb" is a part-of-speech wastebasket, not a semantic class — it collects manner, place, time, frequency, degree, epistemic and connective words with almost nothing in common. Under the two-sort vocabulary, the class splits cleanly, with most members turning out to be procedural.
| Adverb class | Encodes | Type |
|---|---|---|
| Manner (belamente, rapidamente) | a region on a quality dimension, predicated of an event | value |
| Time extent (recentemente) | a region on a temporal-distance dimension | value |
| Frequency, graded (raramente, frequentemente) | a region on the frequency dimension | value |
| Degree / intensity (muito, pouco, bastante) | locate a value against a standard | degree |
| Universal frequency (sempre) | quantification over times | quantification |
| Epistemic, scalar (provavelmente, certamente, possivelmente) | a region on the probability/certainty dimension | value |
| Epistemic, non-scalar (talvez) | epistemic status, no scale named; governs subjunctive | modality |
| Epistemic stance as a concept (dúvida, incerteza, certeza) | the dimension itself | quality |
| Place and deictic time (aqui, hoje, agora) | resolution instruction | reference |
| Conjunctive (contudo, portanto) | discourse relation | connection |
| Affirmation / negation (sim, não, nunca, nada, nenhum) | polarity | polarity |
Manner adverbs are values, not dimensions. Belamente stands to beleza exactly as bonito does — a region on the same dimension, predicated of an event rather than an entity. What the value is predicated of is a frame-level fact, not a type-level one.
Deictics are procedural, not relations. Aqui and hoje encode an
instruction to resolve a referent against the utterance situation — the
textbook case of procedural encoding. Their image-schematic character is
carried by what their frames ground in (FIGURE_GROUND, DEICTIC).
Zero is polarity, not a low value. Nada and nenhum look like values on
the quantity dimension, parallel to pouco. But a low value presupposes
existence — pouco asserts some quantity, just a small one — while nada
and nenhum deny existence outright (¬∃x). The same line separates
raramente (a low but positive frequency, still asserting the event
occurred) from nunca (a negative existential over time points). A zero
reading on any dimension negates that dimension's existential presupposition
rather than measuring a small amount of it. Nada is polarity.
The polarity relation is not the act. Não as propositional polarity is
polarity; the act of affirming or denying (afirmar, negar) is an
event in an @agentive frame. Negação carries both readings under one
form and one POS, and is split into two lemmas accordingly (§1.4).
6.2 Prepositions and adpositions#
ADP is one UPOS covering lemmas on both sides of the conceptual/procedural
line; the split does not follow the tag at all. A preposition may name a
relation, mark a case, express polarity or carry a discourse connection.
| Use | Examples | Type |
|---|---|---|
| Spatial or temporal configuration | sobre, sob, entre, contra, durante, após, dentro de, ao lado de, antes de | relation |
| Privative | sem | polarity |
| Causal, concessive, conditional | por causa de, apesar de, devido a, em vez de | connection |
| Governed by the predicate (case-marking) | de in gostar de, a in obedecer a, em in acreditar em, com in contar com | predication |
| Marking a role the frame already supplies | por (passive agent), com (instrument), para (beneficiary) | predication |
Three diagnostics:
- Contrastive substitution. Holding the governing head constant, can
another preposition replace it with a truth-conditional difference? O
livro está sobre / sob a mesa — yes, so
relation. *Gostar em — no, sopredication. - Questionability. Can the PP answer a wh-question with the preposition inside the answer? Onde? — Sobre a mesa. Governed prepositions fail — there is no question whose answer is de in gostar de.
- Government. Is the preposition listed in the valence of the head that
selects it? If yes,
predicationregardless of what it means elsewhere.
Diagnostic 3 overrides the other two when they conflict: government is observable in the valence data, while the other two rest on judgement.
Governed prepositions are markers, not frame-evoking lemmas. A
case-marking preposition contributes no concept and gets no frame — its
place is the valence pattern, as the marker of an argument position (the role
synsem:marker plays in OntoLex-Lemon). predication records that the lemma
is real and type-bearing but conceptually empty; which head governs it is
recorded at the valence layer, not by a frame link.
Applying the rule. Adpositions are the largest application of §1.4 — one
lemma per (text, POS, type). Splitting stays bounded because readings that
share a type do not split: de has at least six distinguishable readings but
only two types, so it yields two lemmas, not six; genitive, material and
partitive are LUs of the same predication lemma because they are all
markers.
| Lemma | Type | Covers | Examples |
|---|---|---|---|
de₀ |
predication |
government; genitive; material; partitive | gostar de música; o carro de João; mesa de madeira; um copo de água |
de₁ |
relation |
source and origin, spatial or temporal | veio de São Paulo; de segunda a sexta |
a₀ |
predication |
government; dative and recipient marking | obedecer a uma ordem; dei o livro a João |
a₁ |
relation |
goal, direction, temporal location | fui a Roma; às três horas |
em₀ |
predication |
government | acreditar em algo; pensar em você |
em₁ |
relation |
containment, location, temporal location | na mesa; em 2020; no Brasil |
por₀ |
predication |
passive agent; government; price marking | destruído por João; optar por outro; comprei por dez reais |
por₁ |
relation |
path, duration, cause | passou por São Paulo; por três horas; preso por roubo |
com₀ |
predication |
government; instrument and manner marking | contar com você; cortou com a faca; com cuidado |
com₁ |
relation |
comitative, accompaniment | foi com Maria; café com leite |
para₀ |
predication |
government; beneficiary and recipient marking | olhar para a janela; comprei um livro para João |
para₁ |
relation |
goal, direction | fui para Roma; para o norte |
para₂ |
connection |
purpose, linking two events | saiu para comprar pão |
sobre₀ |
relation |
superposition | o livro está sobre a mesa |
sobre₁ |
predication |
topic marking | um livro sobre linguística |
sem₀ |
polarity |
privative — no split, one type covers every reading | sem açúcar; saiu sem pagar |
Sobre splits even though it looks like a paradigm case of a contentful preposition — the split is not a property of "grammatical" adpositions only. Sem does not split at all — the split is driven by type divergence and nothing else.
Disambiguation. Splitting means UPOS filters nothing between de₀ and
de₁ — they share both text and POS. Disambiguation comes from the syntactic
context, using evidence already in the UD parse:
- Government. If the adposition's head is a predicate that selects it —
the
casedependent of anobl/objwhose head lists that preposition in its valence — choose thepredicationlemma. This decides most tokens. - Complement semantics. A place or time complement favours the
relationlemma; an infinitival complement under para forcespara₂. - Default. Where neither fires, prefer the
predicationlemma — the more frequent reading forde,a,em,por,comandpara, and the less damaging error since it attributes no content that is not there.
Extended discussion of the accuracy trade-off this moves to the parser is in Appendix §A.8.
Nothing here is specific to adpositions — the same rule produces
negação₀/negação₁ and pintura₀/pintura₁ in the content lexicon
(§1.4). Adpositions are simply where it fires most often: a small set of very
frequent forms each span the conceptual/procedural line. The split count is
concentrated — a handful of function words account for most of it.
Complex prepositions are MWE lemmas. Dentro de, ao lado de, por
causa de, apesar de take entry_type = mwe and pos = X (§1.2). Their
type follows the table above and is often connection rather than
relation, which is a further reason not to read the type off the tag.
The type cuts across ADP/SCONJ as well as within ADP. Por causa de
(ADP) and porque (SCONJ) are both connection; para + infinitive sits
on the boundary — the tag predicts the type in neither direction.
Consequence for §8. The structural check requires a relation lemma's
frames to have FEs lexicalized by role lemmas. Spatial prepositions
systematically fail this: the FEs of a locative frame are Figure and
Ground, and Portuguese has no role lemmas naming them. The check is
therefore scoped in §8 to relational frames whose FEs have role
lexicalizations.
7. Why coarse: the three layers#
The asymmetry is the most important property of this design; with the type on the lemma it reads as three layers:
| Layer | Grain | Carries |
|---|---|---|
| lemma | coarse identity | the type — what kind of thing the lexeme names |
| LU | sense | the lemma–frame pairing and a sense description — no type |
| frame | situation | the namespace, the participant structure, the grounding schema |
- Destruição is one lemma typed
event. Its LUs may sit in an@agentiveframe (destruction as cause), a@changeframe (destruction as change of state) and a@stativeframe (destruction as resulting condition) — without contradiction and without splitting the lemma. - Aquecer is one lemma typed
event, with LUs inCausar_mudança_de_temperaturaandMudança_de_temperatura. Same type, different frames — the ordinary case. valuebehaves the same way across@attribute,@stativeand@entity. Americano names a region on the origem dimension, full stop, regardless of what it is predicated of: a person (um americano), a thing (um carro americano) or an event (o ataque americano).frm_people_by_originis an@entityframe built overvaluelemmas, and it is not an exception — the live data has dozens more (frm_people_by_religion,frm_people_by_vocation,frm_wealthiness,frm_expertise, and around twenty-five others). This is a construal shift, not a split (§1.4): the lemma staysvalueand the referent varies with what the construction supplies.- POS differences split the lemma, not the LU. Dependenteₙ ("meu
dependente") is
role; dependenteₐ ("ele é dependente") isvalue— two lemmas by §1.2, each with its own type and its own LUs. Unlike americano above, this is not a construal shift: the noun names a fixed institutional role, not "person" as a default referent for the value (§1.4). - Type divergence splits the lemma too. Where one form under one POS carries two types, it is two lemmas by §1.4. Nothing in the LU table records a type, so no divergence can be hidden there.
Construal shifts change neither. Ele é pai shifts a functional concept
to a sortal predicate without changing the underlying concept (Löbner); o
bonito and um americano refer to an entity a value lemma characterizes
without becoming sortals themselves. Neither is type ambiguity, neither
spawns an LU, and neither splits the lemma — the declared type stands, per
commitment 3. §1.4 has the general statement and the diagnostic that tells a
construal shift apart from a genuine split like dependenteₙ/dependenteₐ
above.
8. What this layer validates, and what it cannot#
All checks run on lemma.type_id. There is no second source to reconcile: an
LU's type is its lemma's type, always.
The type licenses a class-compatibility check: does the frame's namespace admit this type at all?
| Namespace | Admits |
|---|---|
@agentive, @change, @phenomenon, @process, @undergoing |
event |
@experiential |
event |
@stative |
event (the maintaining verbs — permanecer, manter — are eventive lemmas like any other verb evoking an eventive frame), quality, value |
@attribute |
quality, value |
@entity |
object, role, value |
@relational |
relation, role |
| — | procedural types are exempt |
A second, structural check:
For every
relationlemma, the frames its LUs evoke should have FEs lexicalized byrolelemmas, and those roles should beisDefinedInthat same frame-as-Relation.
Scope. The check applies only to relational frames whose FEs actually have role lexicalizations. Spatial and temporal frames are excluded: their FEs are Figure and Ground, which no Portuguese lemma names as a role (§6.2).
A relation lemma whose frames have no role lexicalizations is either
mistyped or sitting in a frame missing its participant LUs. A role lemma
whose frames are not Relations is a typing error.
A third check, available only because the type sits on the lemma:
A lemma whose LUs span frames with incompatible admit-sets is a candidate for a split (§1.4).
What the checks do:
- Flag cross-class category errors: a
qualitylemma in an eventive frame, anobjectlemma in a stative frame, arelationlemma in an agentive frame. Automatable, high-precision. Per commitment 3, a flag is a marked combination for review, not a verdict. - Silent on within-class distinctions:
@agentivevs@changevs@phenomenonall admitevent, so the check says nothing about which is right. - Silent on wrong-sense errors: only an annotator working against a real sentence can choose which frame an occurrence belongs to.
- Does not assign types automatically. Which of the six ontological types, or whether a lemma is procedural at all, is annotator judgement guided by §4–§6.
Appendix#
Secondary material: full definitions, theoretical provenance, extended diagnostics, design notes, and open questions. None of it is required to classify a lemma, but §A.1 underlies the criteria in §3.
A.1 Definitions#
Six notions from formal ontology and lexical semantics, used by the diagnostics throughout this document and by the criteria in §3.
Rigidity (Guarino & Welty, OntoClean). A property is rigid if essential to all its instances: anything that is P must be P for as long as it exists (pessoa, cadeira). It is anti-rigid if every instance could cease to be P and survive (comprador, criança, cansado), and semi-rigid if essential for some instances and accidental for others. Anti-rigid properties cannot subsume rigid ones — why pessoa is not a kind of agente. Rigidity is defined here but is not a typing axis (§A.2).
Identity and sortality. A property supplies identity if it carries a
criterion for saying "the same one again"; such properties are sortals,
answering what something is. Pessoa supplies identity; comprador
inherits it from pessoa; vermelho has none. Every object instantiates
exactly one identity-supplying kind. Both object and role are sortals
(§3); what separates them is external dependence, not sortality.
Unity. A property carries unity if all its instances are wholes under one unifying relation (functional, topological, morphological, intentional). Pessoa carries functional unity; água is anti-unity — an arbitrary sub-portion of water is still water. FNBr does not type on unity; it explains why mass nouns need no separate type (§A.3).
External dependence. A property is externally dependent (+D) if every
instance requires the existence of some entity y that is not a part or
constituent of it. Comprador requires a purchase and a seller; filho
requires a parent; cadeira requires nothing. This is the diagnostic that
separates object from role (§3, §4) and replaces constitutivity.
Relationality and uniqueness (Löbner, "Concept Types and Determination", Journal of Semantics 28(3), 2011). A noun concept is relational [+R] if it has an inherent argument slot for a relatum (irmã, filho), and unique [+U] if it determines at most one referent, absolutely or relative to its relatum (sol; pai relative to a child). Crossing the two yields sortal [−R−U], individual [−R+U], relational [+R−U] and functional [+R+U] concepts. FNBr does not type on [U], but both features supply corpus- observable diagnostics (§4) where the ontological test is hard to apply.
Dimension and region (Gärdenfors, Conceptual Spaces; The Geometry of
Meaning). A quality is a dimension — an axis if gradable, a class if
categorial. A region is a value the dimension can take. Temperatura is
a dimension; quente is a region within it — the distinction underlying
quality vs value (§3, §4).
A.2 Why rigidity is not a typing axis#
Rigidity (§A.1) is not used to type anything. Anti-rigid intrinsic
sortals — criança, adulto, novato, defunto, concepts an individual
can stop falling under while continuing to exist, with no second entity
required — are typed object, the same as rigid sortals like pessoa.
The reason is annotator reliability. The dependence test (§4) asks a single existence question: is there some entity, not a part of this one, that must exist? Rigidity asks a modal question: could this individual cease to fall under this concept and still exist? Modal judgements of that kind are unreliable — the same problem that makes the removed state/attribute boundary unreliable (§A.5) — so rigidity is dropped as a typing axis consistently, not as a one-off concession.
Consistent with DUL. dul:Concept ⊑ dul:SocialObject ⊑ dul:Object, so
an anti-rigid intrinsic sortal falls under dul:Object whether or not it is
named separately — a loss of specificity, not an error. It also repairs
role: under a single dependence test, filho is unproblematic even though
being someone's son is arguably rigid (you do not stop being your father's
son when he dies) yet plainly externally dependent.
What is lost. Criança and pessoa carry the same type despite differing
in rigidity, and an export to a rigidity-aware ontology (UFO, gUFO) would
have to decide Kind vs Phase at export time rather than reading it off
the stored type. If that export becomes a requirement, a dedicated phase
type can be reinstated for anti-rigid intrinsic sortals without invalidating
any existing object assignment — a candidate list can be generated by
querying object lemmas whose frames are @entity frames built over
life-stage, career-stage or status dimensions.
A.3 No abstract type#
Abstract nouns need no type of their own. DUL's Abstract is not "abstract
nouns" — it is value spaces: Region, TimeInterval, SpaceRegion, the
space of natural numbers. dul:Object already covers social and cognitive
entities (SocialObject, InformationObject, Concept, Description), so
lei, ideia, teoria, poema are Objects in DUL already.
The only residue is fact and proposition nouns (o fato de que…, a
possibilidade), reified Situations that take event.
Mass nouns need no separate type either: água is anti-unity but a rigid sortal, and unity is not a typing axis (§A.1).
Dot objects (livro = physical • information) are deferred: if
co-predication comes to require it, add a boolean refinement on object —
physical vs social-informational — not a base type (see also §A.9, item 4).
A.4 Grimshaw's nominal diagnostics#
Grimshaw's complex-event/simple-event/result trichotomy separates nominals by argument structure: whether the noun takes obligatory arguments and aspectual modifiers. Simple event nominals (viagem, corrida, evento) fail those tests while remaining ontologically events, so the tests do not identify type (§4). They describe nominal syntax, not this classification.
A.5 quality and value do not distinguish causal origin#
Three reasons back the claim in §4 that causal origin plays no role in the
quality/value classification.
It duplicates the frame. Whether a producing event is on record is a fact
about the situation, which the frame layer already carries (the STATIVE
schema vs Attributes plus a relation to the producing event). A fact that
does not distinguish namespaces, and is already carried by the frame, does
not need a type-level home.
It matches the asymmetry already accepted for destruição. Destruição
stays event while its frames range over @agentive, @change and
@stative; by the same logic quebrado stays value while its frame
carries the fact that the value resulted from an event.
There is no reliable test. A distinction between a "resultant" quality (cansaço, force-maintained postura) and a non-resultant one (beleza, inteligência) — the removed state/attribute line — re-lexicalizes the stage-level/individual-level predicate distinction (Carlson 1977; Kratzer 1995), which is gradient, frequently ambiguous within a single predicate, and which the semantics literature declines to treat as an ontological sort.
Consequences: postura, cansaço → quality. Cansado, quebrado,
sentado → value. Anti-rigidity and causal origin are read off the frame.
A.6 Theoretical grounding for the procedural types#
The conceptual/procedural split is Relevance Theory's (§2). The nine subtypes are not — no RT publication proposes this inventory.
RT subdivides procedural meaning along two axes not used here: the target of the constraint (explicature, higher-level explicature, or implicature — Blakemore 1987; Wilson & Sperber 1993), and the cognitive subsystem the item triggers (Wilson 2011, 2016). The inventory here is organized by functional category instead, applying Escandell-Vidal & Leonetti's thesis that the semantics of functional categories is procedural. Their own enumeration — discourse markers, sentence-mood marks, quotative and evidential particles, intonation, verb tenses and moods, definite determiners and pronouns, deictic and focusing adverbs, information-structure devices — is a list, not a typology; the nine types coarsen that list into classes an annotator can assign.
| Type | Nearest published support | Status |
|---|---|---|
reference |
Wilson & Sperber 1993 on pronouns; Escandell-Vidal & Leonetti on definites and deictics | RT-aligned |
connection |
Blakemore's core case | RT-aligned |
focus |
focusing adverbs in Escandell-Vidal & Leonetti; Iten on concessives and even | RT-aligned |
interaction |
Wharton on interjections; Curcó on Spanish particles; quotative and evidential particles in Escandell-Vidal & Leonetti | RT-aligned |
predication |
the functional-category thesis | derived, not RT |
quantification |
only definite determiners are standardly procedural in RT | departure |
polarity |
classical RT gives não/not a logical entry, i.e. conceptual | departure |
degree |
degree semantics (Kennedy; and Gärdenfors for the scale) | outside RT |
modality |
RT treats only mood and modal particles as procedural | RT-aligned after narrowing (§5) |
quantification and polarity. Classical RT (Sperber & Wilson 1986, ch.
2) treats logical words as concepts with logical entries — a claim about how
an item enters inference. The criterion used here is narrower: does the
lemma name something the conceptualization inventory can hold? Não,
todo and nenhum have no dul: referent under any construal, and typing
them conceptually would require inventing one.
degree is not an RT category. The "region located against a standard"
criterion comes from degree semantics, and the dimension/region framing from
Gärdenfors (§A.1). It is procedural because it instructs rather than names —
the operational test of §2.
modality. Wilson & Sperber (1993) and Ifantidou (1993, 2001) treat
sentence adverbials such as certamente and provavelmente as
conceptual; Papafragou (2000) treats modal verbs as conceptual. In RT
only mood and modal particles are procedural. FNBr's scalar adverbs are
value (§5, "modality vs value"), which keeps this row RT-aligned. One
residue: RT would also call the modal auxiliaries conceptual, following
Papafragou. They are kept procedural because no conceptual type fits them —
they name no scale, no event and no relation — and predication would
misdescribe them as contentless.
What the inventory still lacks. Sentence mood and type (declarative,
interrogative, imperative) has no type of its own — currently spread across
reference (quem) and modality, with free-word mood markers landing in
interaction. Whether mood deserves its own type depends on how much of it
FNBr lexicalizes; mood realized by verb morphology or intonation is not a
lemma and is out of scope. Tense and aspect are procedural on most RT
accounts but out of scope here for the same reason: bound morphology, not
lemmas.
One RT axis considered and declined. Adding target ∈ {explicature, higher-level explicature, implicature} as an orthogonal attribute would let
the scheme cite Wilson & Sperber (1993) directly. It is not adopted: the
target of a procedural constraint varies with use (the same connective can
strengthen or contradict), so the attribute would be a per-token judgement
recorded on a per-lemma row, and nothing in §8 consumes it.
A.7 modality vs interaction for evidentials#
Diz que is interaction on the grounds that it points at an information
source rather than stating speaker confidence. Some frameworks treat
evidentiality as a species of epistemic modality, which would put it with
modality; now that modality is narrowed to non-scalar items, the two are
closer than before. Open — see §A.9, item 11.
A.8 The cost of splitting adpositions#
Splitting moves work from the lexicon to the parser: the lexicon records a stable fact and the parser resolves a contextual one. That is the right direction, but it is a real cost and should be measured on the adposition tokens specifically, not averaged into overall lemma-resolution accuracy — see §A.9, item 9.
A.9 Open questions#
Recorded so they are not rediscovered as objections.
- Validation of
role. What inter-annotator agreement does the dependence test achieve on Portuguese data? The two corpus diagnostics in §4 give an external anchor; the number has not yet been measured. - Rigidity is not recorded. See §A.2.
- [U] is not used. Löbner's uniqueness feature distinguishes filho
(relational) from pai (functional) and predicts determination and
bridging behaviour. FNBr currently collapses both to
role. Whether [U] earns a column is open. - Dot objects.
objectdoes not represent the dual identity of livro (physical • information) that co-predication exposes. Deferred, not solved. - POS × type correlation. How much independent information does the type layer carry over UPOS? One contingency table answers it (§1.5).
- The two-sort split has a boundary. Some lemmas plausibly encode both conceptually and procedurally (sempre, só). The current rule assigns one type; whether that is adequate is untested.
- Split count. How many lemmas split under §1.4, and how concentrated is the count? A few dozen forms, dominated by adpositions, is the expected shape (§6.2). If splits are instead spread thinly across the content lexicon, the type inventory is drawing a line the lexicon does not observe, and the inventory — not the rule — needs revisiting.
- Relatedness between split lemmas is unrecorded. Negação act/operator
and pintura activity/canvas are separate rows with no link. Where the
relation matters it should be added explicitly at the lexical layer
(OntoLex-Lemon
vartrans). Whether it matters enough to build is open. - Adposition disambiguation accuracy. Splitting de, a, em, por, com, para and sobre by type (§6.2) moves their disambiguation to the parser, where UPOS cannot help because the candidates share it. Accuracy on adposition tokens should be measured separately, and the government heuristic evaluated on its own. If it underperforms badly, the alternatives are a richer valence-driven disambiguation or a coarser type inventory for adpositions — not a return to per-LU typing.
- The
modality/valueline needs validating, not deciding. The decision is taken (§5): scalar epistemic adverbs arevalue, talvez and the modal auxiliaries aremodality, on whether the lemma names a lexicalized dimension with a gradable counterpart. What is untested is coverage — whether every PB epistemic adverb falls cleanly on one side. Possivelmente and eventualmente are the likely stress cases, and eventualmente may not be epistemic at all in PB. Run the three tests over a closed list of epistemic adverbs before annotation opens. - The
modality/interactionboundary for evidentials. See §A.7. interactionhas not been tested at scale. It was added for FNBr's pragmatic frames and its three boundaries — againstconnection,modalityandfocus(§5) — are stated but unvalidated. PB interactional particles are frequent and often multifunctional (tá as tag, as agreement token, as imperative of estar), so this type is the most likely source of new §1.4 splits.- The construal-shift diagnostic for substantivized adjectives is new
and untested. The "recoverable covert sortal" test (§1.4) separates
productive ADJ→N ellipsis (o bonito, um americano) from lexicalized
conversion (dependenteₙ, and presumably empregado, acusado and
similar deverbal/deadjectival nouns). Coverage across the lexicon is
unmeasured. The ~25 "people by X"
@entityframes (§7) are themselves a conventionalized middle case — a nationality, religion or vocation adjective defaults to a person referent systematically enough to be frame-licensed, yet no lemma split is made — and it is untested whether other systematic default-referent classes exist that the general test would miss.