Layer 0 — The model in one minute
Happyforce measures continuously, not annually. The architecture has three levels:
Two root indicators, measured directly and continuously:
- Happiness Index (HI) — a proxy for subjective wellbeing. One daily question, four options.
- eNPS — a proxy for commitment. One quarterly question, the standard Net Promoter formula applied to employees.
Six Scores that explain the roots. Each Score decomposes into 3 factors, each factor into 3–4 validated items, all normalized to a common 1–10 scale:
| Score | Factors | Theoretical foundation |
|---|---|---|
| Intrinsic Motivation | Autonomy · Mastery · Purpose | Self-Determination Theory (Deci & Ryan) |
| Relationships | Managers · Peers · Confidence | Gallup meta-analyses; trust-in-leadership research |
| Feedback | Freedom of Opinion · Quality & Frequency · Active Listening | Psychological safety (Edmondson); employee voice |
| Alignment | Values & Ethics · Goals & Role · Trust & Vision | Person-organization fit; goal-setting & role clarity |
| Wellbeing | Stress & Health · Diversity & Equality · Environment | Job Demands-Resources model |
| Reward & Recognition | Compensation · Recognition · Benefits | Organizational justice theory |
The logic that connects them: the Scores are the levers; HI and eNPS are the outcomes. When a Score moves, we can measure how it correlates with the root indicators — at team and segment level, never individual. That is what turns a survey into a driver model: it doesn't just tell you how people are, it tells you which lever is most associated with how they are.
Everything below is the evidence for each piece.
1Why continuous measurement
The dominant instrument in this field is the annual engagement survey. It has a well-documented weakness: it asks people to evaluate in retrospect a whole year of experience, which makes it vulnerable to recency and recall bias — the tendency to answer based on the last few weeks, not the year (Kahneman et al., 2004, developed the Day Reconstruction Method precisely because retrospective evaluation and lived experience diverge).
Continuous measurement takes the opposite approach, closer to experience sampling: many small observations, captured close to the moment they happen. The individual data point is noisier; the resulting time series is far more informative. Three consequences:
- The trend is the signal, not the absolute number. A team at HI 62 tells you little in isolation; a team that dropped from 71 to 62 in six weeks tells you a lot. Our model is built around evolution and deviation, not snapshots.
- You can see events. Annual surveys average events away. Continuous series show them — a leadership change, a restructuring announcement, a policy reversal — with dates attached.
- Response burden collapses. One question a day takes 20–30 seconds. Sustained participation becomes feasible where a 60-item annual survey produces fatigue.
On participation. Responding is voluntary, which introduces self-selection. We mitigate it by (a) monitoring participation as a first-class metric per team and segment, (b) reporting sample sizes alongside every aggregate, and (c) treating shifts in participation itself as a signal — disengagement often shows up as silence before it shows up in scores.
2The Happiness Index (HI)
What it is
One question — "How happy are you at work today?" — with four response options, from "Pretty bad" to "Great". Available every day; each person answers when and if they choose.
How it's computed
Responses are mapped to a 0–100 scale and aggregated with a recency-weighted moving window: recent votes weigh more than older ones. This gives the index two properties by design — it reacts to what's happening now, without whiplashing on a single bad Monday. Aggregates are always computed at team, area or company level (see §6 for anonymity rules).
What it measures — and what it doesn't
HI is a measure of momentary affect, and by accumulation, of subjective wellbeing — the individual's own perception of how they are doing, which is how the field defines wellbeing since Diener's foundational work (1984). This matters: happiness is not something we define for people; it's something each person reports for themselves. The World Happiness Report tradition rests on the same principle — self-reported wellbeing is the measure, not external proxies.
HI is not an engagement score. Mood and engagement are related but distinct constructs; conflating them is a common industry error. In our model, engagement-adjacent constructs live in the Scores (§4); HI is the fast-moving thermometer.
Single-item wellbeing measures have a real research pedigree: they show adequate reliability and strong convergence with multi-item scales (Abdel-Khalek, 2006), and the OECD's Guidelines on Measuring Subjective Well-being (2013) endorse single-item affect measures for exactly this kind of repeated, low-burden measurement. What a single item loses in granularity, it gains in frequency — and frequency is what makes the time series work.
Evidence that it works
Sensitivity to real events. When Spain declared the COVID-19 lockdown in March 2020, HI dropped sharply across our entire client base — an average fall of 19 points within days, visible company by company and team by team. An indicator that did not move under the largest workplace disruption in decades would be measuring nothing. HI moved, immediately and everywhere — and then recovered at different speeds in different organizations, which was itself diagnostic information.
Predictive relationship with turnover. In peer-reviewed research conducted on Happyforce data (Berengueres, Duran & Castro, 2017 — see §7), happiness-derived features contributed to ranking employees by turnover risk. The most interesting finding was methodological: absolute mean happiness was not a significant predictor, while relative happiness — the individual's level normalized by their company's mean — was among the top three. This supports a core design decision of our model: context and deviation carry the information, not raw levels.
3eNPS
What it is
"How likely are you to recommend [company] as a place to work?" on a 0–10 scale. Standard Net Promoter arithmetic (Reichheld, 2003): % promoters (9–10) minus % detractors (0–6), yielding a score from −100 to +100. Measured quarterly.
Why we use it, and its honest limitations
eNPS is coarse — the promoter/detractor cut discards information, and a single evaluative item cannot explain why people would or wouldn't recommend. We keep it for two reasons: it is the most widely understood benchmark in the industry, which makes it useful for boards and cross-company comparison; and as an evaluative measure it complements HI's affective one. In wellbeing research terms: HI captures how work feels day to day; eNPS captures the considered judgment. The two together triangulate better than either alone.
The explanatory work eNPS cannot do is done by the Scores — which is the point of the architecture.
4The Scores model
Architecture
Six Scores → 18 factors → ~57 items. Two item formats, both mapping to a common 1–10 scale:
- Scale items: 1–10 with verbal anchors at both ends (e.g., "My manager brings out the best in us" — Strongly disagree → Strongly agree).
- Option items: 4–5 behaviorally-labeled response options, each mapped to a calibrated score (e.g., "When I ask my manager for help…" — Never (1) / Occasionally (3) / Usually (7) / Every time (10)).
Items are distributed over time rather than delivered as one long questionnaire — the same experience-sampling logic as HI, applied to the driver model. Factor scores aggregate their items; Scores aggregate their factors; everything reports with sample sizes and can be segmented (team, tenure, role, location, or client-defined segments).
Theoretical foundations, score by score
We didn't invent these constructs. Each Score operationalizes a body of organizational psychology with decades of evidence behind it:
Intrinsic Motivation — Autonomy, Mastery, Purpose. This is Self-Determination Theory (Deci & Ryan, 1985; Ryan & Deci, 2000), one of the most replicated frameworks in motivation science: humans are intrinsically motivated when three needs are met — autonomy (control over how you work), competence (growing and using your skills), and purpose/relatedness (your work means something). The Autonomy/Mastery/Purpose formulation was popularized by Pink (2009). Our items map directly: decision latitude and involvement (autonomy), skill development and career visibility (mastery), meaning and contribution (purpose).
Relationships — Managers, Peers, Confidence. The single most robust finding in the engagement literature is that the manager relationship dominates the employee experience: Gallup's meta-analyses across thousands of business units (Harter, Schmidt & Hayes, 2002) consistently place manager-dependent items among the strongest correlates of business outcomes. Peer items draw on the social support and belonging literature; the Confidence factor operationalizes trust in leadership, whose links to performance and attitudes are established meta-analytically (Dirks & Ferrin, 2002).
Feedback — Freedom of Opinion, Quality & Frequency, Active Listening. The Freedom of Opinion factor is a direct operationalization of psychological safety (Edmondson, 1999) — the shared belief that one can speak up without punishment — which two decades of research (including Google's Project Aristotle replication at scale) identify as the strongest team-level predictor of effectiveness. Quality & Frequency and Active Listening draw on the feedback environment and employee voice literatures (Morrison, 2011): voice only exists where someone demonstrably listens.
Alignment — Values & Ethics, Goals & Role, Trust & Vision. Values items operationalize person-organization fit (Kristof, 1996), a consistent predictor of satisfaction, commitment and retention. Goals & Role rests on two classics: goal-setting theory (Locke & Latham) and the role clarity literature (role ambiguity as a chronic stressor — Rizzo, House & Lirtzman, 1970; also the logic behind Gallup's most famous item, "I know what is expected of me at work"). Trust & Vision measures whether strategy is credible from below — a precondition for discretionary effort.
Wellbeing — Stress & Health, Diversity & Equality, Environment. The frame here is the Job Demands-Resources model (Demerouti et al., 2001; Bakker & Demerouti, 2007): strain results from the balance between demands (stress load) and resources (tools, information, support, fair treatment). Stress & Health captures demands and perceived organizational care; Environment captures resources; Diversity & Equality captures inclusion climate (Nishii, 2013), which functions as a resource — exclusion is a chronic demand.
Reward & Recognition — Compensation, Recognition, Benefits. Note what the compensation items ask: not "are you paid a lot" but "are you paid fairly, and can you talk about it". That is organizational justice theory (Adams, 1965; Colquitt, 2001) — perceived fairness of outcomes and processes predicts attitudes and behavior better than absolute pay levels. Recognition items reflect one of the most consistent findings in the engagement canon: regular recognition is a top driver of engagement and retention.
The connecting logic: Scores as drivers of the roots
Because HI, eNPS and Score items are collected continuously from the same population, we can compute — at team/segment level — how each Score, factor, and individual item correlates with the root indicators. This is the analytical core of the model:
- It ranks levers: which dimension is most associated with the wellbeing or commitment of this team, now — not in a generic benchmark.
- It supports the broader evidence base: the engagement→outcomes chain is meta-analytically established (Harter et al., 2002: engagement correlates with productivity, profitability, retention and customer metrics at the business-unit level). Our model localizes that chain to each organization's own data.
On causality. Correlations on observational data identify association, not causation. We treat them as prioritization signals — where to look and act first — and we validate direction the only way observational settings allow: intervene on the lever, and watch the series respond. Continuous measurement is what makes that validation loop possible at all; an annual survey can never close it.
Customization without losing the model
Organizations can adapt the measurement — modify, remove or add items, adjust terminology, define their own metrics — on the same architecture (item → factor → score → root). Custom items follow the same formats and the same 1–10 normalization, so client-specific measurement remains comparable over time and analyzable with the same driver logic.
5From perception data to business metrics
Perception metrics become decision-grade when they connect to costs the organization already recognizes. Our approach, briefly (a fuller treatment lives in our business-case methodology):
- Absenteeism is computed from the client's own hard data (hours lost × salary cost). No estimation involved — this is the conservative floor of any business case.
- Presenteeism — capacity lost while present — is estimated from the soft metrics, always as a range, with every assumption explicit and adjustable by the client. Occupational health research consistently sizes presenteeism costs above absenteeism, so range-based estimates are conservative relative to the academic benchmark.
- Turnover risk builds on the modeling line published in Berengueres et al. (2017): relative and behavioral features, not raw satisfaction levels.
The principle: we provide the model and the calibration from our historical data; the client's own numbers do the monetizing. A business case computed with the client's assumptions is a business case the client believes.
6Anonymity and data ethics
For a measurement system that depends on people telling the truth, anonymity is not a compliance feature — it is a validity condition. People report honestly when honesty is safe.
- Individual responses are never exposed. All reporting is aggregated at team/segment level. Below a minimum group size of 5 people, data automatically aggregates upward to the next level.
- Cross-analysis stays aggregate. Correlation analyses (including any crossing with performance data) run at team/segment level only — never individual. This protects anonymity, participation, and GDPR compliance simultaneously.
- The employee owns the relationship. Each employee accepts Happyforce's terms directly — the data relationship is between the person and Happyforce, not mediated by the employer.
- Infrastructure: EU hosting, GDPR compliance, ISO 27001 certification.
The published research (§7) followed the same standard: the public dataset is fully anonymized, with comment text replaced by character counts.
7Published research and open data
Happyforce is, to our knowledge, one of the few platforms in this category whose data has supported peer-reviewed publication with an openly available dataset:
A collaboration between UAE University and Happyforce, using 2.5 years of anonymized platform data — 34 companies, 4,300+ employees, 221,000+ happiness votes plus anonymous-forum interaction data — to rank employees by turnover risk. Key findings:
- The top three turnover predictors were behavioral and relative, not attitudinal and absolute: likeability (ratio of likes received on one's comments), posting frequency, and relative happiness (individual level normalized by company mean). Precision@50 = 80% on the test set.
- Mean happiness was not a significant predictor. Being at 6/10 in a company averaging 8 is a very different signal from being at 6 in a company averaging 5. Context is the information.
- The anonymized dataset was published openly (Kaggle), allowing independent replication — an unusual level of transparency for commercial workplace data.
These findings fed directly back into the product: it is why our analytics privilege deviation, trend and within-company comparison over absolute benchmarking.
The research line continues on a far larger base:
That accumulated history is what calibrates the models described above.
These figures refer exclusively to organizations measured continuously through the Happyforce platform — the dataset that calibrates the models described in this document. They do not include one-off research initiatives, such as the annual World Happiness at Work survey we conduct with the World Happiness Foundation, whose reach is reported separately.
8Summary of limitations (the section most vendors don't write)
- Voluntary participation → self-selection. Mitigated by monitoring participation as a metric, reporting n everywhere, and reading silence as signal. Not eliminated.
- Single-item roots are coarse. By design — the trade is granularity for frequency. Explanatory depth comes from the Scores, not the roots.
- Correlational driver analysis. Association, not proof of causation. Used for prioritization; validated through intervention and follow-up on the time series.
- Self-report throughout. All perception measures share method variance. This is why the model insists on connecting to hard external data (turnover, absenteeism) rather than living in a closed self-report loop.
- Benchmarks are secondary. Our own research says relative and longitudinal signals beat absolute levels. We show benchmarks because clients ask; we build the model on trends.
References
- Abdel-Khalek, A. M. (2006). Measuring happiness with a single-item scale. Social Behavior and Personality, 34(2).
- Adams, J. S. (1965). Inequity in social exchange. Advances in Experimental Social Psychology, 2.
- Bakker, A. B., & Demerouti, E. (2007). The Job Demands-Resources model: State of the art. Journal of Managerial Psychology, 22(3).
- Berengueres, J., Duran, G., & Castro, D. (2017). Happiness, an inside job? Turnover prediction using employee likeability, engagement and relative happiness. ASONAM '17. DOI: 10.1145/3110025.3110132
- Colquitt, J. A. (2001). On the dimensionality of organizational justice. Journal of Applied Psychology, 86(3).
- Deci, E. L., & Ryan, R. M. (1985). Intrinsic Motivation and Self-Determination in Human Behavior. Plenum.
- Demerouti, E., Bakker, A. B., Nachreiner, F., & Schaufeli, W. B. (2001). The Job Demands-Resources model of burnout. Journal of Applied Psychology, 86(3).
- Diener, E. (1984). Subjective well-being. Psychological Bulletin, 95(3).
- Dirks, K. T., & Ferrin, D. L. (2002). Trust in leadership: Meta-analytic findings. Journal of Applied Psychology, 87(4).
- Edmondson, A. (1999). Psychological safety and learning behavior in work teams. Administrative Science Quarterly, 44(2).
- Harter, J. K., Schmidt, F. L., & Hayes, T. L. (2002). Business-unit-level relationship between employee satisfaction, employee engagement, and business outcomes: A meta-analysis. Journal of Applied Psychology, 87(2).
- Kahneman, D., Krueger, A. B., Schkade, D., Schwarz, N., & Stone, A. A. (2004). A survey method for characterizing daily life experience: The Day Reconstruction Method. Science, 306.
- Kristof, A. L. (1996). Person-organization fit. Personnel Psychology, 49(1).
- Morrison, E. W. (2011). Employee voice behavior. Academy of Management Annals, 5(1).
- Nishii, L. H. (2013). The benefits of climate for inclusion for gender-diverse groups. Academy of Management Journal, 56(6).
- OECD (2013). OECD Guidelines on Measuring Subjective Well-being. OECD Publishing.
- Pink, D. H. (2009). Drive: The Surprising Truth About What Motivates Us. Riverhead.
- Reichheld, F. F. (2003). The one number you need to grow. Harvard Business Review, 81(12).
- Rizzo, J. R., House, R. J., & Lirtzman, S. I. (1970). Role conflict and ambiguity in complex organizations. Administrative Science Quarterly, 15(2).
- Ryan, R. M., & Deci, E. L. (2000). Self-determination theory and the facilitation of intrinsic motivation, social development, and well-being. American Psychologist, 55(1).