It surveys how this term is defined in the published literature — what the sources agree on, where they diverge and what remains contested — and carries its own numbered bibliography. It does not describe what we run. Programme policy is on the core pages.
Our own use of the word is often narrower than general usage; the short definition and the boundary are in the glossary, under Definition of Done. The review date below matters, because the literature moves.
A definitional review. Citations follow IEEE style; see References.
Abstract
Assessment in maker education is the problem of evidencing learning in settings whose defining features — open-endedness, learner-set goals, divergent outcomes — are precisely the features standardised assessment accommodates least well. Survey evidence shows practitioners assess anyway, through self-assessment, portfolios and peer critique rather than tests. This article sets out how the field states the problem, the distinction between process and product evidence, the methods in use, and what the literature does not establish.
I. Definition
The field has no canonical definition; it defines the term by the problem it names. That problem is a claim about artifacts: producing something is not sufficient evidence that anything was learned. Valente and Blikstein put it that “students’ having produced something is not enough to ensure that they have constructed knowledge” [2]. Maker assessment is therefore the practice of establishing what a learner understands, given an artifact that does not by itself demonstrate it.
The practical driver is stated by Rosenheck et al.: fitting the iterative experience of making into school structures is hard partly for want of well-aligned standards and assessments, and of measures for skills such as agency and problem-solving; without a principled way of showing what students know and can do, teachers can neither guide learners nor articulate the value of the curriculum [4].
II. Origins
The conceptual root is Piaget’s distinction between success and understanding, from the two 1974 books translated as Grasp of Consciousness and Success and Understanding. Piaget observed that children use complex actions to achieve premature success — the characteristics of savoir faire — performing a task without understanding how, or being mindful of the concepts involved. Passage from that practical knowledge to understanding runs through the grasp of consciousness, not insight but a level of conceptualisation reached by transforming schemes of action into notions and operations [2].
The empirical root is the portfolio strand: portfolio assessment entered education as a school-based alternative to numeric representations of achievement, and was carried into maker education by Maker Ed’s Open Portfolio Project, whose brief series ran from a 2014 baseline survey through the 2017 survey reported below [1].
III. What distinguishes it
Maker assessment is distinguished by an explicit conflict with the standardised model. Makerspaces are positioned against the “delivered curriculum” and defined learning outcomes attained through standardised assessment; Buxton et al. report the argument that importing that model would “run the risk of squeezing the joy out of learning in a makerspace,” and that standardised assessment culture “struggles to capture the creativity and artistic benefits of makerspaces” [5]. Practitioner data agree: the least prevalent assessment types in and out of school are exactly those most stressed in standard measures — multiple choice (2%), matching items (2%), essay items (6%) — “likely because they’re a poor match to the types of learning occurring in makerspaces” [1].
The discriminating difficulty is not informality but divergence of outcome. Rubrics are far more common in schools (60%) than out of school (7%) because they require a priori planning and tend to stress common outcomes among makers, whereas out-of-school settings typically allow more divergent and emergent ones [1]. The conflict recurs in challenge-based settings, where teachers described open-endedness as conflicting with assessment: testing named concepts is straightforward, but if the goal is “a general way of dealing with challenges, apart from a certain domain, that is very difficult to measure” [7].
IV. Methods in use: what the evidence reports
A. What practitioners actually do
The anchor evidence is a 2017 international survey of 48 makerspaces — 20 in-school, 28 out-of-school [1]. Three-quarters reported assessment measures in place, but split widely: 90% of school-based spaces against 64% of out-of-school spaces. Out of school the leading methods were self-assessment (36%), exit surveys (32%) and peer assessment (29%); in school, self-assessment (65%), rubrics (60%) and portfolios (55%) [1]. The authors read this as evidence that practice is ahead of research: “despite researchers not providing a firm answer on how makerspace learning can be measured, educators in and out of school are moving forward to meet the practical realities” [1].
B. Portfolios and documentation
Portfolios were used by 53% of school-based and 21% of out-of-school makerspaces, and valued differently: 75% of school respondents rated them at least “very important” against 42.8% out of school, leading the report to conclude that portfolio assessment “may not be a one-size-fits-all solution” [1]. The stated rationale is reflective, not credentialing — self-reflection led in both settings (95% in school, 86% out), while college preparation, admissions and career development were the least-mentioned reasons out of school [1]. Respondents disagreed that documentation takes time from making or interrupts flow; the reported barriers were access to dedicated technology (23%), privacy and consent (14.5%) and lack of youth motivation to capture work (14.5%). Portfolio implementation correlated more with youth documentation practices than with staff practices, suggesting such systems become sustainable only when learners drive the capture [1].
C. Process versus product evidence
This is the field’s central methodological distinction. In product-based assessment a rubric is applied to the final artifact; in process-based assessment educators and peers track development and apply criteria to intermediary milestones [3].
Blikstein and Valente argue product-based assessment has specific unintended consequences. Pursuing a better final result, students split labour by existing ability, “which ends up reinforcing inequalities in the classroom” — the experienced take technical tasks, the less experienced manual ones. A good product can obfuscate a poor process or very little learning; conversely one that looks unfinished or non-functional “can be the result of a very rich learning experience — especially for novices with ambitious goals” [3]. A second argument concerns drift: premade objects and ever more capable machines make it “easy to produce objects that are relatively sophisticated without any deep understanding of how they work” — making an LED blink took an engineering degree and tens of hours in the 1980s and takes minutes with a modern kit today — so if the standard is the production of objects, assessment stands on “increasingly shaky ground” and the makerspace risks becoming “a mere production facility of cool and curious contraptions” [3].
The empirical counterpart is Rosenheck et al.’s finding that a process skill such as the patience to iterate “is neither visible on final product-oriented assessments, nor in any single observation or reflection” — it emerges only from evidence gathered over time [4].
D. Formative and embedded assessment
The dominant formative design is embedded assessment, integrating assessment into curricular materials so assessment and learning happen in tandem. The Beyond Rubrics Toolkit studied by Rosenheck et al. runs in three phases: setting context, evidence collection (artifacts that make process skills visible during the work rather than after) and meaning making (joint interpretation of the body of evidence) [4]. From coding that evidence the authors derive six qualities of a strong body of evidence — alignment, action, specificity, articulation, abstraction and coherence across tools and sessions — presented as “a temperature check” rather than a threshold, coherence being a property of the whole collection rather than any single artifact [4]. Formative assessment is not thereby made easy: Zhang notes the counter-argument that it “is elusive as it has to depend on a teachers’ experience and intuition to a great extent,” much of it addressing teachable moments that arise unpredictably [6].
E. Self- and peer-assessment
Self-assessment is the single most common method across surveyed makerspaces (48% overall), and a third of sites scaffold it with sentence starters — “I had difficulty when…,” “I solved my challenge by…,” “Did you use a new tool? Which one? How was it used to make your project?” Peer assessment, defined as critique or guided comments by a fellow participant, was in use at 33% of sites [1].
The dependency this creates is explicit: Beyond Rubrics “relies on students’ ability to reflect on their experience and assess their own learning. Without the ability to do this, they will not be able to generate evidence of learning” [4]. Self-assessment grants autonomy in place of a passive role, but requires psychological safety, teacher modelling and active support for accurate responses, and is a skill developed over time rather than assumed [4]. It can also fail on contact: Buxton et al. report a pilot in which children asked to assess their own learning often did not, citing lack of time and the observation process being a “mood killer” [5].
F. Observation frameworks and standardised instruments
Where rubrics are rejected, structured observation often replaces them. The Makerspace Learning Assessment Framework adapts the Characteristics of Effective Learning used in the English Early Years Foundation Stage — chosen because they “describe how children learn rather than what they learn” — into five observed areas (playing and exploring; active learning; critical thinking; creativity and design; social learning), each with six prompts applied as an observation schedule alongside narrative field notes, on an explicit “process not product” rationale [5]. A distinct, psychometric route also exists: Shen et al. validated a 52-item self-report maker literacy scale across four dimensions — learning, practical, interdisciplinary and innovation ability — using Delphi expert review, factor analysis and item response theory on 4,983 secondary students across 24 Chinese provinces [10]. This buys measurement properties at the cost of assessing reported disposition rather than observed work.
G. The interview method: eliciting understanding of one’s own artifact
The most theoretically grounded method is Piaget’s clinical-critical method, applied to maker settings by Valente and Blikstein [2]. It “consists in the systematic intervention by an educator based on the learner’s conduct, such as verbal interaction, the manipulation of objects, or an explanation.” The educator presents a problem situation that is conceptually rich and significant to the learner, observes what the learner does, and intervenes to clarify the meaning of those actions or explanations — forming a hypothesis about an action’s meaning and attempting to confirm it through the intervention itself [2].
The authors flag one characteristic as “generally not described in the studies regarding this theme,” or even in Piaget’s own work: the educator’s prior examination of the activity in terms of the concepts involved and their different levels of complexity. Without that preparation the conversation cannot locate the learner’s level of conceptualisation [2]. Their worked example is a catapult: the teacher asks the student to explain how it works, how the structure was developed, whether the launch distance or angle can be changed, and how the spring’s tension affects the distance travelled [2]. Crucially the method is not only diagnostic — the educator’s questions and challenges, pitched within the learner’s zone of proximal development, “contribute to the process of elaborating new conceptual relations” [2]. It assesses and teaches in the same move.
Two supporting strategies accompany it. Analysis of product testing: for a test to contribute to knowledge construction it must make explicit the variables observed, the procedures used and the analysis of the data, after which the teacher can question the student on conclusions and improvement [2]. And where programmable or digital fabrication tools are involved, the commands given to the machine constitute an action representation — the concepts and strategies used, made inspectable and debuggable — which “can be seen as a ‘window into the mind’ of the learner” [2]. That window is narrower than it appears: in their authors’ response, Blikstein and Valente warn that a learner can compensate for missing hardware knowledge in software and vice versa with the same functional result, so the lens “has to be embedded in carefully designed activities, process-based assessments, and project milestones” [3].
V. Scope of adoption
Assessment is widespread in maker settings but weakly connected to decisions: 33% of respondents said it informed instructional design, 16% future programming, 8% funding and administrative decisions, 8% professional development, and 36 of the 48 sites planned to increase portfolio assessment [1].
Reflection — the ingredient most of these methods depend on — is not reliably present. A systematic review of library makerspace research reports a study finding a makerspace was perceived as encouraging nearly all types of innovation behaviour and exploration but not reflection, with the authors recommending that questioning and reflection be deliberately encouraged [9].
VI. Limitations
The evidence base is thin relative to the practice. The anchor survey is 48 self-selecting sites, and its own longitudinal comparison rests on two independent samples with different respondents [1]. The observation-framework study states that its findings “cannot be extrapolated to all children, as this was not an experimental study” [5].
The central distinction is argued, not measured. The preference for process over product evidence rests on arguments from consequence — labour division, obfuscation, technological drift — rather than demonstrated superior validity [3]. None of the sources reviewed here reports predictive or criterion validity for a maker assessment; the only formal validation evidence is the internal reliability and factor structure of a self-report scale [10].
The methods assume a reflective, articulate learner. Embedded assessment depends on students’ capacity to describe their own learning [4], and self-assessment is the most common method in use [1]. Every inference is bounded by the quality of evidence learners can produce — and the six-qualities study offers no threshold for when a body of evidence is good enough to interpret, only a heuristic [4].
Assessment of the individual can misattribute the environment. Vossoughi et al. observe that empirical studies of making tend to foreground individual learning processes rather than joint activity or explicit analyses of teaching. Narrating one recorded episode twice, they show a learner’s success is readily attributed to the learner, while a learner’s struggle becomes “a story of [his] struggles rather than an examination of the pedagogical supports and relationships available in the environment” [8]. A record capturing only the individual will read the presence or absence of support as a property of the person.
Little establishes what any of it is for. The starkest finding in the survey is the gap between near-universal adoption of assessment and its very limited influence on programming, funding or staff development [1] — assessment is performed in maker settings well ahead of any settled account of what it should be used to decide.
References
[1] K. Peppler, A. Keune, F. Xia, and S. Chang, “Survey of Assessment in Makerspaces,” Open Portfolio Project Research Brief 17, Maker Ed, 2018.
[2] J. A. Valente and P. Blikstein, “Maker Education: Where Is the Knowledge Construction?” Constructivist Foundations, vol. 14, no. 3, pp. 252–262, 2019.
[3] P. Blikstein and J. A. Valente, “Professional Development and Policymaking in Maker Education: Old Dilemmas and Familiar Risks,” Authors’ Response, Constructivist Foundations, vol. 14, no. 3, pp. 268–271, 2019.
[4] L. Rosenheck, G. C. Lin, R. Nigam, P. Nori, and Y. J. Kim, “Not all evidence is created equal: assessment artifacts in maker education,” Information and Learning Sciences, 2021, doi: 10.1108/ILS-08-2020-0205.
[5] A. Buxton, L. Kay, and B. Nutbrown, “Developing a Makerspace Learning and Assessment Framework,” in Proc. 6th FabLearn Europe / MakeEd Conf. 2022, Copenhagen, Denmark, 2022, doi: 10.1145/3535227.3535232.
[6] D. Zhang, “STEM Education through Making: What Are Affordances and Challenges of Making Out of School Club?” Open Journal of Social Sciences, vol. 9, no. 9, 2021, doi: 10.4236/jss.2021.99042.
[7] K. Helker, M. Bruns, I. M. M. J. Reymen, and J. D. Vermunt, “A framework for capturing student learning in challenge-based learning,” Active Learning in Higher Education, vol. 26, no. 1, pp. 213–229, 2024, doi: 10.1177/14697874241230459.
[8] S. Vossoughi, P. K. Hooper, and M. Escudé, “Making Through the Lens of Culture and Power: Toward Transformative Visions for Educational Equity,” Harvard Educational Review, vol. 86, no. 2, pp. 206–232, 2016.
[9] S. H. Kim, Y. J. Jung, and G. W. Choi, “A systematic review of library makerspaces research,” Library and Information Science Research, vol. 44, no. 4, 101202, 2022.
[10] G. Shen, J. Huang, and G. Wang, “Development and validation of a maker literacy assessment scale for secondary school students in China,” Humanities and Social Sciences Communications, 2026, doi: 10.1057/s41599-026-07804-w.