
Version 2.2.0
Note on this version: This guide consolidates the two previous SDK user guides — BSocial SDK v3.2.0 (general-purpose / social robotics, kiosks, mobile) and BAutomotive SDK v2.1.0 (in-cabin automotive) — into a single unified BSocial product. Where a feature originated in one SDK only, it is labelled (Automotive) or (Core) so you can tell at a glance what applies to your deployment. Please review and confirm feature availability and licensing tiers before publishing, as this merge was done from the source documents only.
BSocial is an innovative SDK for automatic behaviour data processing and analysis. It has been designed and developed to be low computational resource demanding, easy to use and work with standard colour, as well as NIR images.
BSocial can detect, track, and process up to 5 people.
BSocial can be deployed on an edge device, that is a low power and low cost device, which makes it a great fit for integration into consumer products, such as robots, cars, mobile device applications and so on. The SDK does not require a network connection to operate or to send data to a remote server for processing as it is capable of performing input data processing and analysis locally, which additionally makes it a great fit for products which might not have a stable networking connection. BSocial is extremely portable and as such does not require any special hardware, such as a specific GPU device(s) or specialised imaging hardware in order to function properly.
The SDK is written entirely in C++; however, to facilitate potential integration with another product, as well as general ease of use, it comes with a variety of APIs, which include C++, C, Python and Unity C#. Each of these interfaces can be used to communicate with BSocial to send input data into it and obtain the resulting predictions back.
BSocial is extremely flexible and can be deployed either standalone or as a part of another software product, including as an in-cabin monitoring system for vehicles. (Automotive) BSocial features CAN bus support and is capable of sending resulting data directly into CAN to be used by other onboard in-car systems — see CAN Bus Integration below.
BSocial is a cross-platform SDK supporting a variety of operating systems and CPU architectures. The purely C++ nature of implementation BSocial features allows for re-use of a single code base to support, and if needed to expand, a whole range of software and hardware BSocial is compatible with. As such, the SDK currently comes in a number of different platform specific variations, each of which is presented below along with the corresponding platform specific requirements:
linux-x64 version, which is meant for deployment under Linux on a device featuring an x86_64 CPU architecture. This version of the SDK requires Ubuntu 20.04 / Debian 11 or newer, or an alternative Linux environment featuring GLIBC of version 2.31 or newer.linux-arm64 version, which is meant for deployment under Linux on a device featuring an arm64 CPU architecture. This version of the SDK requires Ubuntu 20.04 / Debian 11 or newer, or an alternative Linux environment featuring GLIBC of version 2.31 or newer.BSocial variations targeting Linux systems currently additionally require a copy of LIBGTK-3.0 installed system-wide in order to run the SDK command line tool, however this is not a requirement to run the SDK itself.
darwin-x64 version, which is meant for deployment under macOS on an Apple desktop or notebook computer featuring an x86_64 CPU architecture. This version of the SDK requires macOS 10.15 or newer.darwin-arm64 version, which is meant for deployment under macOS on an Apple desktop or notebook computer featuring an arm64 Apple Silicone CPU architecture. This version of the SDK requires macOS 12.0 or newer.ios-arm64 version, which is meant for deployment under iOS on Apple mobile devices. This version of the SDK requires iOS 13.0 or newer.android-arm64-v8a version, which is meant for deployment under Android on mobile devices, as well as robotic systems. This version of the SDK requires Android 8.0 or newer.windows-x64 version, which is meant for deployment under Windows on a desktop PC or laptop computer featuring a x86_64 CPU architecture. This version of the SDK requires Windows 10 or newer.All of the BSocial variations listed in the previous section come with a Python API supporting Python versions 3.7.x to 3.12.x, with the exception of the darwin-arm64 and linux-arm64 versions, which lack support for Python 3.7.x, however do support Python versions 3.8.x to 3.12.x.
Variations of the SDK targeting Apple and Windows computers are primarily maintained for additional ease of the SDK presentation and evaluation and as such might lack certain functionality elements otherwise present in the Linux versions of BSocial.
(Automotive) macOS builds of BSocial currently lack CAN bus support.
This section presents a brief overview of the BSocial functionality. For more in-depth information on the matter please refer to the Technical Application Guide and BSocial Functionality Overview document.
BSocial SDK offers a wide range of functionality on automatic behaviour analysis. These functionality elements include:
Once started, BSocial first tries to detect a face in the input images fed into the API. If it succeeds BSocial then tries to perform facial landmarks tracker initialisation using the face detection bounding box as the initialisation area. If it fails to locate the landmarks from the facial bounding box it does not attempt to perform any further operations on the input image since all of them require a valid set of facial landmarks. Once BSocial manages to initialise the landmarks tracker it continues to track the landmarks skipping the face detection step as it is only needed to initialise the tracker whenever it loses track of the facial landmarks.
Please note that in order to ensure the optimal quality of the resulting predictions, the SDK enforces a minimum face size check and will reject faces of insufficient resolution. The minimum face size required is currently set to 141 pixels along the diagonal.
The maximum number of faces the SDK is currently able to track and re-identify is limited to 5. This number can be limited further if needed using the set_max_faces_to_track API method.
BSocial can also re-identify people who have previously presented themselves on camera or video processed with the SDK. This allows to associate behaviour of a specific person with previous behavioural data of the same person. For example, a person can leave in a bad mood, and when they return, BSocial can identify if they are feeling better. In combination with multi-person processing, this also allows to clearly distinguish which behaviour data belongs to which individual in the scene.
In order to perform re-identification, BSocial generates and assigns every person a unique face ID. The SDK is capable of saving generated face IDs in order to perform persistent re-identification, that is a re-identification across multiple sessions, each being a camera feed or a video file.
By default, persistent user re-identification is off. It can be enabled on a user by user basis by the process of enrollment of neutral face images, which can be achieved either with the command line tool or via the SDK API. The relevant command line tool arguments are:
--enroll <NEUTRAL_IMAGE_PATH>--list-enrolled-users--remove-all-enrolled--remove-enrolled-user <USER_NAME>--remove-enrolled-id <USER_ID>The relevant API methods are:
get_neutral_facesclear_neutral_facesadd_neutral_faceadd_neutral_face_to_existing_identityremove_neutral_face_with_persistent_face_idremove_neutral_face_with_user_defined_namePersistent user re-identification data is stored in locations, which differ depending on the host system platform:
$HOME/Library/Application Support/BM.BSocial$HOME/.local/share/BM.BSocial%USERPROFILE%\AppData\Roaming\BM.BSocialThis location can be adjusted via set_neutral_data_directory method of the SDK API.
Re-identification is only possible at the moment if a person's face is in relatively frontal view of the camera and if the input image quality is sufficient as recognised by the SDK. The re-identification functionality is not intended to be used for authentication.
A number of internal components of BSocial currently adopt temporal awareness to further enhance quality of the predictions produced by the SDK. The API provides interface to feed images into BSocial one by one, however in order to comply with the temporal awareness functionality of the internal components in BSocial it is expected that these images are consecutive elements of the same continuous image sequence.
The API includes a method called reset() introduced to be able to clear the state of the internal temporal variables of BSocial. This way a single instance of the API can be used to process multiple videos while ensuring correct handling of the temporal information. It is therefore recommended to call the API's reset() method before processing of a new video is started as this will ensure temporal information accumulated by BSocial over one video will not have any effect on the following videos should a single instance of the BSocial API be used to process multiple videos.
The majority of predictions BSocial generates are produced on the per-frame basis. One exception is the expressed mood estimation, which requires the SDK to process a number of video frames before it can produce a valid expressed mood prediction. The most common way of obtaining a valid expressed mood prediction is to call the retrieve_mood() method of the SDK API at the end of video processing to generate a single mood score summarising the mood level over the entire video. Please see the Python API Reference and C++ API Reference sections below for details on how to use the retrieve_mood() method to obtain a mood score.
In the context of BSocial, Confusion is defined as a facial expression characterised by lowering of eye-brows and squinting of eyes. The SDK measures apparent expressed confusion which may or may not always correlate with felt/experienced confusion.
The SDK calculates and outputs so-called ImHO confidence score, which represents an overall prediction confidence reported by BSocial. This score ranges between 0 and 1 and is calculated based on a number of factors calculated by the SDK, which include Visibility and Frontal Pose scores.
Visibility score ranges from 0 to 1 and is reflecting how well a given tracked face is visible on camera. For instance if the input image is too dark this will drive Visibility score down.
Frontal Pose score also ranges from 0 to 1 and is reflecting how acceptable head rotation of a given tracked face is with respect to the camera. If head position is far from optimal for the SDK to perform accurate analysis this score will get lower, driving the overall ImHO score down as well.
Interaction engagement is defined as a user's behavioural focus and intent to interact over time. BSocial calculates a continuous engagement score characterized by physical markers such as sustained head and gaze alignment, eye aspect ratio (alertness), physical proximity (leaning in), and conversational cues (voice activity). BSocial measures apparent behavioural engagement, which may or may not always correlate with the user's internal cognitive state or genuine interest.
BSocial can estimate whether a tracked subject is showing signs of drowsiness over time, producing a drowsiness state flag and a confidence score. See the Drowsiness CSV output section and the BSocialDrowsiness structure in the API Outputs Reference for details.
BSocial can detect coughing and sneezing events from a tracked subject, each with an associated confidence score. See the Sneeze Cough CSV output section and the BSocialSneeze structure in the API Outputs Reference for details.
In in-vehicle deployments, BSocial can transmit resulting predictions directly onto a CAN bus for consumption by other onboard in-car systems, using the init_canbus() and send_canbus_message() API methods. See the Python API Reference and C++ API Reference sections below for details. This capability currently requires a Linux target and is not available on macOS builds.
BSocial requires a licence key in order to operate. Licence keys are generated on the per-client basis and are created to cover, if applicable, a specific time period of the SDK operation and in some cases specific functionality that is made available to a client. Any limitations imposed on the licence key come from a prior contractual usage agreement between the client and Blueskeye. A copy of the licence key is normally shipped to the client along with a copy of the SDK.
The SDK supports both online and offline licence verification through a unified mechanism. Licence keys may require an Internet connection to perform online verification upon BSocial initialisation via a request to a remote licence verification server. Alternatively, if an Internet connection is not required, the key is verified offline. The type of verification is determined based on the licence key provided.
A licence key can be passed into BSocial as a string or alternatively as a licence file with the .bskai extension containing the licence key. Each key is a string of characters of variable length. To learn more about how to pass a licence key into the SDK as a string, please see the set_licence_key method description in the API reference section(s) of this user guide. To learn more about how to pass a licence key into the SDK with a licence file, please see the load_licence_key method description in the API reference section(s) of this user guide. Please note that depending on code security regulation applied to certain builds of the SDK the load_licence_key method might not be available in the API.
By default all copies of BSocial are shipped with online keys unless the end solution BSocial is meant to be integrated into explicitly requires offline licence verification.
The SDK supports multiple tiers of data output, controlled dynamically by the capabilities encoded in licence keys. These output levels determine which data points are returned via the C++ API structures and which columns are written to the CSV output files.
If a licence key does not specify an output level, the SDK will default to the Standard level.
The Standard level provides a robust set of core facial and emotional data points, optimised for general use cases. This is the default output level for evaluation licences and standard client integrations.
Included Data (with applied precision rounding):
Please note that Action Units and advanced Gaze Origin data are excluded from this tier.
The Advanced level is provided to trusted partners who require deep access to granular facial muscle movements and raw origin data. It includes all features of the Standard level, plus the following additions:
Additional Data (with applied precision rounding):
BSocial SDK supports standard colour, as well as NIR input image processing. In both cases as of now it requires images of 8-bit long unsigned integer pixel type in the range between 0 and 255. Images supported by the SDK can be one of the following types:
Image dimensions can vary, however the minimum face size shown in the input image is expected to be of at least 256 by 256 pixels in size, otherwise a reduction in the output processed data quality is expected to occur.
BSocial features built-in image quality assessment and as such performs automatic evaluation of every image passed into the SDK to account for various environmental factors, which could affect input image quality. Images are evaluated in terms of their brightness and blurriness. Images which are found to be too dark, too bright or too blurry are reported via the SDK API as any of these factors is expected to result in the loss of the output processed data quality.
The SDK supports a number of different camera placement options inside a vehicle, whether it is a left-hand-drive or right-hand-drive car. These placement options as of now include:
See also the Technical Application Guide's Hardware Requirements appendix for detailed camera specifications.
In addition to visual input data, certain features of BSocial rely on input audio processing. The SDK supports audio capture devices capable of capturing audio data of the following format characteristics:
BSocial SDK is portable and as such does not require any specific installation procedure. In order to get a copy of the SDK simply download an archive containing one of the variations of BSocial matching a supported platform of choice:
BM.BSocial-<SDK_VERSION>-linux-x64.zip, which targets linux-x64 platform.BM.BSocial-<SDK_VERSION>-linux-arm64.zip, which targets linux-arm64 platform.BM.BSocial-<SDK_VERSION>-darwin-x64.zip, which targets darwin-x64 platform.BM.BSocial-<SDK_VERSION>-darwin-arm64.zip, which targets darwin-arm64 platform.BM.BSocial-<SDK_VERSION>-windows-x64.zip, which targets windows-x64 platform.BM.BSocial-<SDK_VERSION>-ios-arm64.zip, which targets ios-arm64 platform.BM.BSocial-<SDK_VERSION>-android-arm64-v8a.zip, which targets android-arm64-v8a platform.For a more in-depth overview of the various platforms supported by BSocial please see the System Requirements section above.
If no download link is made available please contact a member of the Blueskeye team to request a copy of the SDK, as well as a licence key file to go with it.
Once a copy of the SDK is downloaded simply unpack it to an arbitrary location on disk. Location of the uncompressed copy of BSocial will be referenced to as <SDK_PACKAGE_PATH> or simply as BSocial SDK package path throughout the rest of this user guide.
BSocial comes packaged with everything needed for its correct operation with the exception of the licence key, which comes outside of the SDK package. Every copy of BSocial comes with the following set of files and folders, which are listed below in alphabetical order:
3rdparty_licenses - folder containing licence files of all 3rd-party open-source projects the SDK depends on and was compiled with.bin - folder containing the BSocial command line tool called BM.BSocial-app, which enables standalone usage of BSocial, as well as BM.BSocial-tests, which provides an easy way to quickly test the BSocial installation on a chosen set of hardware.etc - folder containing sample image(s), video(s), as well as other types of files, which are used by the BSocial command line tool to test the SDK installation.include - folder containing the SDK API headers. These header files enable BSocial linking with another software project.lib - folder containing compiled BSocial shared library file(s). These files contain the SDK implementation. On macOS these files have .dylib extension, on Linux these files have .so extension and on Windows there is a pair of .lib and .dll files.python - folder containing compiled binaries of the SDK Python interfaces, one per each of the major supported versions of Python. Please see the Python Usage Examples section below for detailed instructions on how to use the Python interface of BSocial presented via an example usage scenario. Please note that this folder is missing in the Windows version of the SDK. The Python binding binaries on Windows can be found in the bin folder.SDK User Guide.pdf - a copy of this user guide in PDF format.SDK User Guide.md - a copy of this user guide in plain text format.BSocial offers a quick and easy way to test whether the SDK functions properly on a chosen set of hardware. It is highly recommended to perform the post-installation tests covered in this section once a copy of the SDK is downloaded and unpacked. These instructions are universal and apply to any platform specific variation of BSocial, and whether the testing is performed under Linux, macOS or Windows.
To run the set of the post-installation tests BSocial SDK offers, open a terminal emulation application on host device, navigate to the bin folder of the downloaded SDK package and execute the following command:
./BM.BSocial-tests -l <LICENCE_KEY_PATH>
where <LICENCE_KEY_PATH> is a placeholder indicating location of the licence key file.
The command specified above launches the SDK post-installation testing routine. It normally takes only a few seconds to perform all of the included tests. The following message is then should be printed to the console indicating that all tests have been completed successfully:
[----------] Global test environment tear-down
[==========] 66 tests from 4 test suites ran. (107520 ms total)
[ PASSED ] 60 tests.
In case of any technical difficulties during the post-installation testing stage of BSocial, as well as at any other point while using the SDK please feel free to contact a member of the BlueSkeye team for support at help@blueskeye.com.
In addition to the tests executable covered in the previous section, BSocial SDK also comes bundled with the command line tool called BM.BSocial-app, which can be used to perform data analysis without any additional software.
BM.BSocial-app is capable of performing analysis of video and audio files, as well as live feeds from cameras and microphones. It can also optionally visualise the analysis and save the resulting predictions to a CSV file.
The SDK command line tool offers its own comprehensive help, which can be printed to console as follows:
./BM.BSocial-app -h
The help screen lists all of the arguments the tool supports, along with a brief description of what each of these arguments allows to achieve. There are only two arguments which are required, and if not specified BM.BSocial-app is configured to explicitly ask for them. These are license key file path and a source of input video, which can be a camera or a video file. The rest of the arguments are optional and can be provided as needed. For example:
--log-level <LOG_LEVEL> argument can be specified to control the amount of console output the tool and the SDK generate.--csv-* group of arguments can be specified to enable additional columns in the output CSV files.--print-devices argument can be specified to get a list of available camera and microphone devices the tool can connect to.--print-machine-id argument can be specified to print a unique machine ID generated by BSocial for the host system. This ID can be used to produce offline license key tied to a specific computer.The command line tool can be manually interrupted by pressing the CTRL+C key combination, which gracefully terminates analysis. If the tool is run with visualisation enabled at the end of video processing the UI window can be closed by double pressing Q key on keyboard.
Please note, that not all audio and video file formats and codecs are currently supported by the BSocial command line tool. The exact set of supported formats and codecs depends on which platform BSocial is built for. At the moment Windows and Linux builds of BSocial come with a set of precompiled FFMpeg shared libraries linked with the SDK command line tool, which allow the command line tool to process video formats supported by FFMpeg version 4.2.2 compiled with LGPLv2.1 license setting. On macOS the SDK command line tool can work with all of the codecs supported by the Apple AVFoundation framework, which comes included with every version of macOS supported by BSocial.
Please also note that on macOS the tool will need to request your permission to access camera and / or microphone if you choose to connect to one of them. Once permission is granted the tool will need to be restarted. This will only need to be done once as macOS acknowledges access permissions on per executable level.
A known current limitation of the tool is its inability to work with audio data stored within a video file. To process audio data along with video data the tool currently requires audio path specified via a separate argument -a in .wav format.
To make it easier to get started with BSocial, this section introduces several typical usage examples of the SDK command line tool presented below:
Analyse input video file:
./BM.BSocial-app -l <LICENCE_KEY_PATH> -v <VIDEO_PATH>
Analyse live camera images and visualise analysis using audio device at index 0 and camera device at index 0:
./BM.BSocial-app -l <LICENCE_KEY_PATH> -c 0 -m 0 -d
Analyse live camera images using camera device at index 0:
./BM.BSocial-app -l <LICENCE_KEY_PATH> -c 0
Analyse live camera images, visualise analysis and save the resulting processed data to a CSV file using audio device at index 0 and camera device at index 0:
./BM.BSocial-app -l <LICENCE_KEY_PATH> -c 0 -m 0 -d -o <OUTPUT_CSV_PATH>
Analyse live camera images using audio device at index 0 and camera device at index 0:
./BM.BSocial-app -l <LICENCE_KEY_PATH> -c 0 -m 0
Analyse a pair of input video and audio files and visualise analysis:
./BM.BSocial-app -l <LICENCE_KEY_PATH> -v <VIDEO_PATH> -a <AUDIO_PATH> -d
BSocial allows to save the resulting processed data to a CSV file.
To reduce file sizes and speed up data processing, the SDK uses a default configuration that only saves the most frequently used columns. All other columns are considered optional and must be explicitly enabled.
The default CSV output includes:
CSV file composition can be adjusted by passing an instance of the BMCSVConfig structure to the start_write() API method, which allows to toggle individual sections of the CSV on or off. Please note that if a section that is restricted by license level is requested, the SDK will omit it and print a warning to the console.
CSV output can be enabled either by passing an output path to the SDK command line tool or by enabling this functionality via the BSocial API. The way the output CSV can be generated with the command line tool has already been covered in the previous sections. The same output CSV can also be generated via the SDK API with the following methods the API offers: set_output_path(), start_write() and stop_write(). The first method allows to pass an output path to the SDK API to specify where it should write the resulting processed data to. The other two methods allow to specify at which point BSocial should start writing to the output CSV file and at which point it should stop respectively. A typical way to use these methods is to call set_output_path() and start_write() before an input video processing starts and then call stop_write() once the video processing completes. This way a single instance of the BSocial API can be used to process multiple videos and create a new CSV file for each of the videos analysed with the SDK. For more in-depth information about these API methods please see the Python API Reference and the C++ API Reference sections below.
The following subsections describe each available CSV output block and whether they are included by default or optional.
The output CSV file BSocial creates represents a table, where one or more rows correspond to each video frame and every column corresponds to a piece of the resulting processed data the SDK generates. For each video frame BSocial creates a minimum of one row in the output CSV file. If there is no face detected in the given video frame that row is populated with zeros. If there is only face detected in the given video frame that row is populated with the predictions generated for that face detection. If there is more than one face detected in the given frame the SDK writes a row per each of detections. The first row of the CSV contains labels of each of the columns to ensure all of the data BSocial generates is clearly labelled. These data is organised into a number of groups, each corresponding to specific functionality element the SDK offers. These groups include:
FrameFace IDFace DetectionHead PositionHead OrientationFace LandmarksGazeAction UnitsAffectImHOFace Image QualityDrowsiness (Automotive)Emotion ZoneVVADAVAD (Core)Sneeze Cough (Automotive)AttentionConfusion (Core)Interaction Engagement (Core)Most of these groups include a dedicated Valid column, such as Face Detection Valid for example. These are so-called data validity columns, values of which can be either 0 or 1. Please note that if a validity value for a group is different from 1 values in the rest of the columns forming the corresponding group should for a given frame be ignored.
An overview of each of these resulting processed data groups is presented next in this section.
Default Output: Yes (Included by default)
The Frame group of the output data is dedicated to store video frames information. It includes the following columns populated by BSocial for every video frame passed into the SDK:
Frame Index (zero-based)Frame Width (in pixels)Frame Height (in pixels)Default Output: No (Must be explicitly enabled)
The Face ID group of the output data is dedicated to store face re-identification results. Face re-identification functionality helps BSocial manage its internal temporal variables buffer and ensure predictions accumulated while tracking one person do not affect predictions generated for another person should there be more than one face in a video sequence processed with the SDK. It includes the following columns populated by BSocial for every video frame passed into the SDK:
Face ID ValidFace ID CurrentFace ID CountThe first column Face ID Valid contains 1 if BSocial was able to successfully re-identify a face in a given frame and 0 otherwise. The Face ID Current column contains a numerical ID value of the person currently being tracked. The Face ID Count column contains the total number of faces BSocial was able to identify in the current video sequence.
Default Output: No (Must be explicitly enabled)
The Face Detection group of the output data is dedicated to store face detection results. It includes the following columns populated by BSocial for every video frame passed into the SDK:
Face Detection ValidFace Detection XFace Detection YFace Detection WidthFace Detection HeightThe first column Face Detection Valid contains 1 if BSocial was able to successfully detect a face in a given frame and 0 otherwise. The Face Detection X and Face Detection Y columns contain X and Y pixel coordinates of the top left corner of the detected face bounding box. The last Face Detection Width and Face Detection Height columns contain width and height of the face bounding box in pixels.
Default Output: No (Must be explicitly enabled)
The Head Position group of the output data is dedicated to store head position estimation results. It includes the following columns populated by BSocial for every video frame passed into the SDK:
Head Position ValidHead Position XHead Position YHead Position ZThe first column Head Position Valid contains 1 if BSocial was able to successfully estimate head position in a given frame and 0 otherwise. The Head Position X and Head Position Y columns contain X and Y pixel coordinates of the centre of the head. The Head Position Z column contains an estimation of relative distance between the head and the camera.
Default Output: No (Must be explicitly enabled)
The Head Orientation group of the output data is dedicated to store 3D head orientation estimation results. It includes the following columns populated by BSocial for every video frame passed into the SDK:
Head Orientation ValidHead Orientation YawHead Orientation PitchHead Orientation RollThe first column Head Orientation Valid contains 1 if BSocial was able to successfully estimate 3D head orientation in a given frame and 0 otherwise. The Head Orientation Yaw, Head Orientation Pitch and Head Orientation Roll columns contain the corresponding head orientation angle values in degrees.
Default Output: No (Must be explicitly enabled)
The Face Landmarks group of the output data is dedicated to store tracking results of the 68 different facial points. It includes the following columns populated by BSocial for every video frame passed into the SDK:
Face Landmarks ValidFace Landmarks X0 ... Face Landmarks X67Face Landmarks Y0 ... Face Landmarks Y67The first column Face Landmarks Valid contains 1 if BSocial was able to successfully track the 68 facial points in a given frame and 0 otherwise. The rest of the columns contain X and Y pixel coordinates of each of the facial points indexed from 0 to 67.
Default Output: No (Must be explicitly enabled)
The Gaze group of the output data is dedicated to store social gaze direction tracking results. It includes the following columns populated by BSocial for every video frame passed into the SDK:
Gaze ValidGaze Angle YawGaze Angle PitchGaze Angle Mean YawGaze Angle Mean PitchGaze Vector XGaze Vector YGaze Vector ZGaze Vector Mean XGaze Vector Mean YGaze Vector Mean ZGaze Left Eye OccludedGaze Right Eye OccludedGaze Left Eye TrackedGaze Right Eye TrackedGaze Head Pose FallbackThe first column Gaze Valid contains 1 if BSocial was able to successfully track social gaze direction in a given frame and 0 otherwise. The following two columns Gaze Angle Yaw and Gaze Angle Pitch contain the corresponding social gaze direction angle values in degrees. The next two columns Gaze Angle Mean Yaw and Gaze Angle Mean Pitch contain the same angles, but smoothed over a short time period. The Gaze Vector X, Gaze Vector Y and Gaze Vector Z columns contain the same social gaze direction data presented as a 3D unit vector, rather than angles in degrees. The last three columns Gaze Vector Mean X, Gaze Vector Mean Y and Gaze Vector Mean Z also contain the social gaze direction 3D unit vector, but smoothed over a short period of time. The Gaze Left Eye Occluded and Gaze Right Eye Occluded columns denote if the eyes are occluded or closed.
The last three columns Gaze Left Eye Tracked, Gaze Right Eye Tracked and Gaze Head Pose Fallback identify which eye(s) were tracked in a given frame and if the SDK had to fall back to head pose while attempting to resolve gaze direction respectively.
Default Output: No (Must be explicitly enabled)
The Action Units group of the output data is dedicated to store intensity estimation results of the 15 FACS Action Units the SDK currently supports. It includes the following columns populated by BSocial for every video frame passed into the SDK:
Action Units ValidAction Units ConfidenceAction Units AU01, AU02, AU04, AU05, AU06, AU07, AU09, AU10, AU12, AU14, AU15, AU17, AU23, AU25, AU45The first column Action Units Valid contains 1 if BSocial was able to successfully estimate Action Unit intensities in a given frame and 0 otherwise. The second column Action Units Confidence contains a floating point value indicating how confident the SDK is about the given set of Action Unit intensity estimation predictions, ranging from 0 to 1. The rest of the columns contain intensities ranging from 0 to 5 for each of the supported Action Units.
Default Output: Yes (Included by default)
The Affect group of the output data is dedicated to store dimensional affect Valence, Arousal and Dominance estimation results. It includes the following columns populated by BSocial for every video frame passed into the SDK:
Affect ValidAffect ValenceAffect ArousalAffect DominanceAffect Valence ConfidenceAffect Arousal ConfidenceAffect Dominance ConfidenceThe first column Affect Valid contains 1 if BSocial was able to successfully estimate Valence, Arousal and Dominance levels in a given frame and 0 otherwise. The following three columns Affect Valence, Affect Arousal and Affect Dominance contain estimated Valence, Arousal and Dominance levels respectively, ranging from -1 to 1. The last three columns contain the corresponding prediction confidence values.
Default Output: Yes (Included by default)
The ImHO group of the output data is dedicated to store overall confidence scores calculated by the SDK on a per-frame basis based on the assessment of input data qualities performed by BSocial. It includes the following columns populated by BSocial for every video frame passed into the SDK:
ImHO ValidImHO ScoreImHO Visibility ScoreImHO Frontal Pose ScoreThe first column ImHO Valid indicates validity of the calculated scores. The following three columns contain values of the corresponding scores ranging from 0 to 1.
Default Output: No (Must be explicitly enabled)
The Face Image Quality group of the output data is dedicated to store automatic input image quality assessment results. It includes the following columns populated by BSocial for every video frame passed into the SDK:
Face Image Quality ValidFace Image Quality AcceptableFace Image Quality BlurryFace Image Quality Too DarkFace Image Quality Too BrightFace Image Quality NIREach of these columns contain a value of either 0 or 1. The first column Face Image Quality Valid contains 1 if BSocial was able to successfully perform the image quality assessment and 0 otherwise. The second column Face Image Quality Acceptable contains 1 if the input image quality was found acceptable and 0 otherwise. The following columns contain 1 if an image was found to be one of the following: too blurry, too dark or too bright. The last column Face Image Quality NIR contains 1 if the image was identified as NIR and 0 otherwise.
Default Output: No (Must be explicitly enabled)
The Drowsiness group of the output data is dedicated to store subject drowsiness estimation results over time. It includes the following columns populated by BSocial for every video frame passed into the SDK:
Drowsiness ValidDrowsiness StateDrowsiness ConfidenceThe first column Drowsiness Valid contains 1 if BSocial was able to successfully estimate the presence or absence of drowsiness in a given frame and 0 otherwise. The next column Drowsiness State contains a value of either 0 or 1, where 0 stands for the absence of drowsiness and 1 stands for the presence of drowsiness. The last column Drowsiness Confidence contains the drowsiness presence confidence value.
Default Output: Yes (Included by default)
The Emotion Zones group of the output data is dedicated to store the subject's emotion zones classification results. The list of emotion zones currently supported by the SDK includes:
AngryEcstaticHappyHorrifiedNeutralSadShockedSurprisedThe Emotion Zones group of the output data includes the following columns populated by BSocial for every video frame passed into the SDK:
Emotion Zones ValidEmotion Zones ConfidenceEmotion ZoneEmotion Zone Angry ProbabilityEmotion Zone Ecstatic ProbabilityEmotion Zone Happy ProbabilityEmotion Zone Horrified ProbabilityEmotion Zone Neutral ProbabilityEmotion Zone Sad ProbabilityEmotion Zone Shocked ProbabilityEmotion Zone Surprised ProbabilityThe first column Emotion Zones Valid contains 1 if BSocial was able to perform a successful classification of emotion zones and 0 otherwise. The next column Emotion Zones Confidence holds prediction confidence of emotion zones estimation. The next column Emotion Zone contains a string label of the emotion zone classified as the closest representation of the subject's present emotional state, that is the emotion zone with the highest classification probability score. The rest of the columns contain floating point values ranging from 0 to 1 indicating classification probability scores of each of the emotion zones supported by the SDK.
Emotion zones classification depends on Valence and Arousal predictions and as such requires a small temporal input buffer in order to generate a valid prediction. This results in a small delay while processing input video frames before BSocial starts to output valid emotion zones classification results.
Default Output: No (Must be explicitly enabled)
The VVAD group of the output data is dedicated to store the subject's visual voice activity detection results. It includes the following columns populated by the SDK for every video frame passed into the SDK:
VVAD ValidVVAD StateVVAD ConfidenceThe first column VVAD Valid contains 1 if BSocial was able to successfully detect presence or absence of visual voice activity in a given frame and 0 otherwise. The next column VVAD State contains a value of either 0 or 1, where 0 stands for the absence of visual voice activity and 1 stands for the presence of visual voice activity. The next column VVAD Confidence is a value between 0 and 1 which denotes the confidence of the model that the subject is talking where 0 is not talking and 1 is talking.
Default Output: No (Must be explicitly enabled)
The AVAD group of the output data is dedicated to store the subject's audio voice activity detection results. It includes the following columns populated by the SDK for every video frame passed into the SDK:
AVAD ValidAVAD StateAVAD ConfidenceThe first column AVAD Valid contains 1 if BSocial was able to successfully detect presence or absence of audio voice activity in a given frame and 0 otherwise. The next column AVAD State contains a value of either 0 or 1, where 0 stands for the absence of audio voice activity and 1 stands for the presence of audio voice activity. The next column AVAD Confidence is a value between 0 and 1 which denotes the confidence of the model that the subject is talking where 0 is not talking and 1 is talking.
Default Output: No (Must be explicitly enabled)
The Sneeze Cough group of the output data is dedicated to store sneezes and coughs detection results. It includes the following columns populated by BSocial for every video frame passed into the SDK:
Sneeze Cough ValidSneeze Cough Detected CoughSneeze Cough Detected SneezeSneeze Cough Confidence NegativeSneeze Cough Confidence CoughSneeze Cough Confidence SneezeThe first column Sneeze Cough Valid contains 1 if BSocial was able to successfully detect presence or absence of either coughs or sneezes in a given frame and 0 otherwise. The following two columns contain a value of either 0 or 1, where 0 stands for absence and 1 stands for presence of coughs and sneezes respectively. The last three columns contain the corresponding cough and sneeze detection confidence values.
Default Output: No (Must be explicitly enabled)
The Attention group of the output data is dedicated to store the subject's gaze-to-object mapping results. It includes the following columns populated by the SDK for every video frame passed into the SDK:
Attention ValidAttention Looked AheadAttention Looked ObjectAttention Gaze Origin X, Y, ZAttention Gaze Vector X, Y, ZAttention Projected Intersection X, YThe first column Attention Valid contains 1 if BSocial was able to successfully map gaze to a specific object or area in a given frame and 0 otherwise. The next column Attention Looked Ahead contains a value of either 0 or 1, where 0 stands for the looking ahead state and 1 otherwise. The third column Attention Looked Object contains a label of the object the subject appears to be looking at, where applicable.
The following columns contain 3D coordinates of gaze origin, gaze vector direction, as well as 2D coordinates of the point marking intersection of the gaze vector with the attention area of interest (AOI) in the AOI coordinate space.
Default Output: No (Must be explicitly enabled)
The Confusion group of the output data is dedicated to store the confusion estimation score. It includes the following columns populated by BSocial for every video frame passed into the SDK:
Confusion ValidConfusion ScoreThe first column Confusion Valid contains 1 if BSocial was able to successfully assign a confusion score to the face being tracked and 0 otherwise. The second column Confusion Score contains the actual confusion score, which can range from 0 to 1, where 0 indicates not being confused and 1 indicates being confused.
Default Output: No (Must be explicitly enabled)
The Interaction Engagement group of the output data is dedicated to store the user's behavioral engagement and interaction intent results. It includes the following columns populated by BSocial for every video frame passed into the SDK:
Interaction Engagement ValidInteraction Engagement ScoreIntention to InteractThe first column Interaction Engagement Valid contains 1 if BSocial was able to successfully calculate the engagement metrics in a given frame and 0 otherwise. The following column Intention to Interact contains a value of either 0 or 1, where 0 stands for the absence and 1 stands for the presence of a user's intent to interact. The Interaction Engagement Score column contains the final computed continuous engagement confidence value.
In addition to the command line tool covered in the prior sections, the BSocial SDK can also be used in a Python script via the SDK Python interface. This section presents two example usage scenarios of the BSocial Python interface. The first scenario demonstrates the most basic usage example showing how to initialise an instance of the SDK, load just a single sample image, run inference on the loaded sample image and obtain the resulting processed data. The second, more involved scenario, demonstrates how the basic usage example can be extended to process a video with BSocial in Python.
Additional information on the BSocial Python interface, including the complete list of available Python API function calls can be found in the Python API Reference section below.
Please ensure the version of Python installed on the host computer is supported by BSocial. As of now BSocial's Python interface supports Python versions 3.7.x to 3.12.x, with the exception of the macOS Apple Silicone, as well as Linux ARM64 host computers where the Python version support is limited to Python 3.8.x to 3.12.x.
Python version can be verified by executing the following command in a terminal emulation application:
python --version
Please note that on some platforms the executable python points to a Python 2.x.x version of the interpreter. On such platforms typically a Python 3.x.x interpreter is called python3.
In order to start using BSocial in Python, first add the path to the folder containing the compiled BSocial Python interface binaries to the list of Python system paths. This is the path called python located within the uncompressed BSocial package. This can normally be done as follows:
import sys
sys.path.insert(0, "<SDK_PACKAGE_PATH>/python")
where <SDK_PACKAGE_PATH> is a placeholder indicating location of the uncompressed BSocial package on disk.
Once the python interface binaries path is added to the list of Python system paths, the next step is to import the BSocial Python module:
from BMBSocial import BMBSocialAPI, BSocialImageType
Should at this point Python report the following error at the SDK Python module import stage please ensure you are running a supported version of Python:
Traceback (most recent call last):
File "example.py", line 10, in <module>
from BMBSocial import BMBSocialAPI
ImportError: No module named BMBSocial
Upon successful import of the BSocial Python module, create a new instance of the SDK API:
api = BMBSocialAPI()
and load a licence key:
assert api.load_licence_key("<LICENCE_KEY_PATH>") == 0
or, alternatively pass the license key directly into the SDK:
api.set_licence_key("<LICENCE_KEY>")
Once licence verification checks are complete, the next step is to initialise the instance of the BSocial API:
assert api.init() == 0
At this point BSocial is ready to analyse input data and it is now possible to pass a sample image into the instance of SDK API. For example, this can be done as follows using a widely popular OpenCV library:
import cv2
im = cv2.imread(<SAMPLE_IMAGE_PATH>)
api.set_image(im, BSocialImageType.BGR, False)
In the example code snippet above <SAMPLE_IMAGE_PATH> refers to the path of the sample image on disk.
Please note that for the purpose of simplicity of this specific example only a single image is passed into the SDK API. In reality however once the SDK API is initialised, this and the following steps should be repeated for every image in a continuous image sequence as BSocial is specifically designed to work with continuous image sequences, such as video files, rather than standalone images. If the end goal is to process multiple continuous image sequences with a single instance of BSocial, the API's reset() method must be called before switching from one sequence to another. More details about the reset() method of the API can be found in the Python API Reference, as well as C++ API Reference sections below in this user guide.
Once the image is passed into the instance of the BSocial API the next step is to run the analysis:
api.run()
and retrieve the image processing results upon completion of the analysis:
predictions = api.get_predictions()
predictions is a list of objects populated by the get_predictions method with all available data generated by the SDK for the last image passed into the API. An overview of the fields included in the predictions structure can be found in the API Outputs Reference section of this user guide. The list contains an entry per each face detected by the SDK. For the purposes of this example, the code snippet below illustrates a way a set of Valence and Arousal predictions generated for the last image passed into the API can be obtained from the populated predictions structure provided the SDK was able to detect at least a single face in the input image:
if len(predictions) > 0:
if predictions[0].affect.valid:
valence = predictions[0].affect.valence
arousal = predictions[0].affect.arousal
This section of the user guide demonstrates how the basic usage scenario covered above can be expanded to process a video with the BSocial SDK. Please note, that this example requires the Python package of OpenCV installed in your Python environment.
First, import the Python interface of the SDK and initialise BSocial API as shown below following the sequence of steps already covered in the previous section:
import sys
sys.path.insert(0, "<SDK_PACKAGE_PATH>/python")
from BMBSocial import BMBSocialAPI
from BMBSocial import BSocialImageType
api = BMBSocialAPI()
assert api.load_licence_key("<LICENCE_KEY_PATH>") == 0
assert api.init() == 0
Once the API is successfully initialised, the next step is to open a video file with OpenCV and retrieve the FPS value of the video:
import cv2
capture = cv2.VideoCapture("<VIDEO_PATH>")
fps = capture.get(cv2.CAP_PROP_FPS)
Once a new video is ready to be processed with BSocial, the next step is to reset the API:
api.reset()
Next, pass the video FPS value into the API:
api.set_video_fps(fps)
It is also possible to make BSocial save predictions generated while processing the video into a CSV file. This can be achieved with the following API method calls, the first of which allows to specify the path to the resulting CSV file, while the second allows to tell BSocial to start writing to the file path specified:
api.set_output_path("<OUTPUT_CSV_FILEPATH>")
api.start_write()
Alternatively, to enable optional columns, instantiate and pass a configuration object instead:
config = bsk.BMCSVConfig()
config.includeGaze = True
config.includeActionUnits = True
api.start_write(config)
The next step is to iterate over video frames as shown in the example code below. First, a variable called im is populated with image data using the OpenCV method call read(), then this variable is passed into BSocial with the set_image() method call. Once the image is passed into the API, the following run() and get_predictions() method calls are used to process the image and retrieve the resulting predictions respectively.
Finally, the overlay() method of the BSocial SDK can optionally be called to draw generated prediction visualisations on top of the input image and the follow-up OpenCV function calls imshow() and waitKey() can be called to display the input image overlaid with the predictions generated on screen.
while True:
success, im = capture.read()
if not success:
break
api.set_image(im, BSocialImageType.BGR, False)
api.run()
predictions = api.get_predictions()
if len(predictions) > 0:
if predictions[0].affect.valid:
valence = predictions[0].affect.valence
arousal = predictions[0].affect.arousal
api.overlay(im, BSocialImageType.BGR, False)
cv2.imshow("window", im)
cv2.waitKey(10)
Once video processing is complete, call api.stop_write() method to tell the BSocial SDK to stop writing to the output CSV file:
api.stop_write()
In addition to the BSocial command line tool covered in the previous section, BSocial can also be used in a C++ program via the SDK C++ API header. This section presents an example usage scenario of the BSocial C++ interface. It demonstrates how to initialise an instance of the SDK, pass input data into an instance of the SDK, run inference and obtain the resulting processed data.
The first step in a C++ program is to include the BSocialAPI.hpp C++ API header and create a new instance of the SDK API:
#include <BSocialAPI.hpp>
bsk::BSocialAPI *mBSocialAPI = new bsk::BSocialAPI();
The BSocial C++ headers can be found within the unpacked SDK in the include folder.
Once the instance of the SDK API is created, the next step is to load a licence key:
std::string licence_key_path = "<LICENCE_KEY_PATH>";
int rcode = mBSocialAPI->load_licence_key(licence_key_path);
if (rcode != 0) {
return rcode;
}
or, alternatively if load_licence_key method is not available in a particular build of the SDK, pass the license key directly into the SDK using the set_licence_key method as shown below:
std::string licence_key = "<LICENCE_KEY>";
mBSocialAPI->set_licence_key(licence_key_path);
and initialise the newly created instance of the BSocial API:
rcode = mBSocialAPI->init();
if (rcode != 0) {
return rcode;
}
All BSocial C++ methods returning a numerical return code, such as the init() method demonstrated in the code snippet above, return 0 on success and 1 or a greater positive value otherwise.
Once the API is initialised, it is recommended to run the API's reset method, which should be called to clear the internal SDK state in order to correctly re-use the same API instance to process multiple videos:
mBSocialAPI->reset();
Optionally, call the set_min_process_time method to set the minimum amount of time for BSocial to spend processing a single frame of input video. This method takes a floating point argument defining the amount of such time in milliseconds. This can be helpful in scenarios where the amount of CPU resource consumed by the SDK should be limited. In case the SDK finishes processing a frame faster than the minimum amount of time set, BSocial will sleep before returning from the run method thus freeing CPU resource for other software running in parallel. The following example code sets the minimum processing time to 100 ms. The default value is 66 ms. The SDK can be made to run as fast as possible by passing 0 as the minimum processing time.
mBSocialAPI->set_min_process_time(100);
Next, specify the framerate of the input video using the set_video_fps method. In the example code below it is set to 30 FPS:
mBSocialAPI->set_video_fps(30);
At this point BSocial is ready to analyse input data and it is now possible to start passing input audio and video data into the SDK. First, prepare and set input audio data:
const int audio_sample_rate = 44100;
const int audio_channels = 1;
const int audio_format_length = 2;
unsigned char* audio_buffer;
const int audio_buffer_length;
mBSocialAPI->set_audio(
audio_buffer,
audio_buffer_length,
audio_sample_rate,
audio_channels,
audio_format_length
);
The code snippet above defines the following set of variables needed to correctly pass audio data into an instance of the BSocial API:
audio_buffer is a pointer to a byte array taken by the set_audio API method, which contains input audio data. The length of this array may vary depending on the amount of time passed between set_audio calls. Once passed into the set_audio method call, the contents of this array will be copied by BSocial and stored within the SDK instance for internal audio buffering. The caller is responsible for allocating this array, populating it with audio data and releasing the array after it is passed into the SDK with the set_audio method call.audio_buffer_length is an integer taken by the set_audio API method, which defines the length of the input audio buffer passed into the set_audio method. This number may vary depending on the amount of time passed between set_audio calls. The caller is responsible for setting this variable to a correct value before passing it into the set_audio method.audio_sample_rate is an integer taken by the set_audio API method, which defines the sample rate of the audio data passed into the SDK. Sample rate is measured in Hz. It defines the number of samples in the input audio data corresponding to 1 second of audio. Please note that at the moment the SDK only supports a sample rate of 44100 Hz and will not accept audio data of a sample rate different from 44100.audio_channels is an integer taken by the set_audio API method, which defines the number of audio channels in the audio data passed into the SDK. Please note that at the moment the SDK only supports single channel mono audio data and will not accept audio data with a channel count greater than 1.audio_format_length is an integer taken by the set_audio API method, which defines the length of the audio format passed into the SDK. Format length is measured in bytes. Please note that at the moment the SDK only supports audio data of signed 16-bit long integer type, which is 2 bytes long, and as such will not accept values of format length different from 2.Once audio data is passed into the SDK, the next step is to prepare and pass the input video frame:
const int img_width;
const int img_height;
unsigned char* img_data;
mBSocialAPI->set_image(
img_width, img_height, img_data, bsk::BSocialImageType::BGR);
The code snippet above defines the following set of variables needed to correctly pass image data into an instance of the BSocial API:
img_width is an integer taken by the set_image API method, which defines the width of the input image in pixels.img_height is an integer taken by the set_image API method, which defines the height of the input image in pixels.img_data is a pointer to a byte array containing pixel data in unsigned 8-bit long format.All of these variables need to be initialised before they are passed into the set_image method. The last argument taken by the set_image method specifies the type of the image passed into the SDK as one of the types defined by BSocialImageType. In the example above it is set to BGR, which is typical when using the OpenCV library to load and process images and videos. Please see the set_image method description in the API reference section(s) below for more information on the arguments this method takes and the complete list of supported image types.
Please note that for the purpose of simplicity of this specific example only a single image is passed into the SDK API. In reality however once the SDK API is initialised, the set_audio and set_image methods, as well as all of the following steps, should be repeated for every image in a continuous image sequence as BSocial is specifically designed to work with continuous audio-visual sequences, such as video files, rather than standalone images.
Once the image and the audio are passed into the instance of the BSocial API the next step is to run the analysis:
mBSocialAPI->run();
and retrieve the resulting predictions upon completion of the analysis:
std::vector<BSocialPredictions> predictions;
mBSocialAPI->get_predictions(predictions);
predictions is a vector of structures of type BSocialPredictions, which gets populated by the get_predictions method with all available data generated by the SDK corresponding to the last input data pair passed into the API. An overview of the fields included in the predictions structure can be found in the API Outputs Reference section of this user guide. The list contains an entry per each face detected by the SDK. For the purposes of this example, the code snippet below illustrates a way a set of Valence and Arousal predictions corresponding to the last input data pair passed into the API can be obtained from the populated predictions structure provided the SDK was able to detect at least a single face in the input image:
if (predictions.size() > 0)
{
if (predictions[0].affect.valid)
{
float valence = predictions[0].affect.valence;
float arousal = predictions[0].affect.arousal;
}
}
Once the instance of the BSocial API is no longer needed it can be disposed of as follows:
delete mBSocialAPI;
Additional information on the BSocial C++ interface, including the complete list of available C++ API function calls, can be found in the C++ API Reference section below.
BSocial is a fairly straightforward SDK to link against for C++ development. There is only one public API header and one shared library to link against. The public API header can be found under the include folder of the SDK package; it is called BSocialAPI.hpp. The shared library can be found under the lib folder of the SDK package; it has .dylib extension on macOS, .so extension on Linux and .dll extension on Windows. There could be additional shared libraries present under the lib folder, however there is no need to link against them. An example linking procedure with another C++ project using CMake is presented below:
set(BSK_SDK_PATH "<SDK_PACKAGE_PATH>")
add_library(BSK_SDK SHARED IMPORTED)
if(APPLE)
set_target_properties(BSK_SDK PROPERTIES IMPORTED_LOCATION
${BSK_SDK_PATH}/lib/libBM.BSocial-lib.dylib)
elseif(MSVC)
set_target_properties(BSK_SDK PROPERTIES IMPORTED_LOCATION
${BSK_SDK_PATH}/lib/libBM.BSocial-lib.dll)
else()
set_target_properties(BSK_SDK PROPERTIES IMPORTED_LOCATION
${BSK_SDK_PATH}/lib/libBM.BSocial-lib.so)
endif()
target_link_libraries(${PROJECT_NAME} BSK_SDK)
target_include_directories(${PROJECT_NAME} PRIVATE "${BSK_SDK_PATH}/include")
Method called to retrieve the SDK version string. This method must be called once the API constructor is called. Takes no arguments.
BMBSocialAPI.version() -> str
Method called to load a licence key into an instance of the BSocial API. This method must be called once per every created instance of the API after the API constructor is called, but before any other API methods. Takes a single string argument licence_key_path, which is used to specify the path to the licence key file. This method returns 0 on success and 1 otherwise.
BMBSocialAPI.load_licence_key(licence_key_path) -> int
An alternative method called to pass a licence key into an instance of the BSocial API. This method must be called once per every created instance of the API after the API constructor is called, but before any other API methods. Takes a single string argument licence_key containing the licence key.
BMBSocialAPI.set_licence_key(licence_key) -> None
Method called to set the logging level. By default the SDK logs info, warning and errors. This can be changed so that the logging is less verbose by setting the log level to a lower value.
BMBSocialAPI.set_log_level(log_level) -> None
Method which can be called to limit the maximum number of faces tracked by the SDK. This method must be called before the SDK is initialised. Takes a single integer argument indicating the desired limit of faces the SDK should track. This number must be greater or equal to 1 and smaller or equal to 5. Returns 0 on success and 1 otherwise if the maximum number of faces requested is outside of the allowed range.
BMBSocialAPI.set_max_faces_to_track(max_faces_to_track) -> int
Method called to enable or disable incremental inference. By default incremental inference is enabled. Incremental inference allows to noticeably reduce the amount of CPU resource utilised by the SDK by preventing it from running inference of certain models every time the SDK's run method is called. Instead, the SDK runs inference of these models on every other frame, while retaining the last successful prediction until the next inference takes place. Takes a single boolean argument, which can be set to True in order to enable incremental inference or False to disable it.
BMBSocialAPI.set_inference_increment_enabled(enabled) -> None
Method called to set the minimum amount of time the SDK spends processing a single frame. By default, the limit is set to 66 ms, which is equivalent to 15 frames per second. This method can be used to adjust the amount of compute resource BSocial consumes per second in order to make it available to other software running alongside the SDK, if applicable. Takes a single floating point argument time in milliseconds. If a value of 0 is supplied the SDK will run as fast as possible on the given hardware with no processing time limitations. If the minimum frame processing time is set and exceeded a warning will be printed to the console.
BMBSocialAPI.set_min_process_time(time) -> None
Method called to initialise the SDK API. This method must be called once per every created instance of the API after the SDK API constructor and one of the licence key setting methods described above are called, but before any of the following API methods. Takes no arguments. This method returns 0 on success, 1 on the SDK initialisation failure and 2 on license key verification failure.
BMBSocialAPI.init() -> int
Optional SDK API method which can be used to reset the internal state of the API instance, which is useful when switching from processed video to another to reset the SDK internal temporal variables. Takes no arguments.
Please note that this method also resets the output CSV file path set with the set_output_path method, as well as the target video FPS set with the set_video_fps method, in addition to the SDK internal temporal variables. This method should therefore be called before both set_output_path and set_video_fps.
BMBSocialAPI.reset() -> None
Method called to load a set of object mapping coordinates used for gaze-to-object mapping. This method must be called once the SDK API instance is initialised. Takes a single string argument AOI_positions_path specifying the path to the object(s) mapping data file. Returns 0 on success and 1 otherwise.
BMBSocialAPI.load_object_mappings(AOI_positions_path) -> int
Method called to load a set of camera calibration parameters used for gaze-to-object mapping. This method must be called once the SDK API instance is initialised. Takes a single string argument camera_calibration_path specifying the path to the camera calibration parameters data file. Returns 0 on success and 1 otherwise.
BMBSocialAPI.load_camera_calibration(camera_calibration_path) -> int
Method which can be optionally called to set a custom set of gaze angles corresponding to the subject observed looking forward position. By default both of the looked ahead gaze angles are zero. This method must be called once the SDK API instance is initialised. Takes two float arguments pitch and yaw, which represent the corresponding gaze angles.
BMBSocialAPI.set_look_ahead_angles(pitch, yaw) -> None
Method which can be optionally called to set a custom set of camera rotation and translation parameters to indicate the location of the camera in space for the purposes of gaze-to-object mapping. By default all of these parameters are set to zero. This method must be called once the SDK API instance is initialised. Takes six float arguments x, y, z, roll, pitch and yaw, which represent the camera position and orientation respectively.
BMBSocialAPI.set_camera_transform(x, y, z, roll, pitch, yaw) -> None
Optional SDK API method which can be used to initialise BSocial's internal CAN subsystem to enable predictions the SDK generates to be transferred via CAN bus. This method must be called once the SDK API instance is initialised. Takes two arguments. The first argument is a string specifying the target CAN device name, e.g. vcan0. The second argument is an integer specifying the start index of the messages the CAN subsystem generates. Returns 0 on success and 1 otherwise.
BMBSocialAPI.init_canbus(deviceID, startIndex) -> int
Optional SDK API method which can be used to send a packet of data through BSocial's internal CAN subsystem. This method must be called once the SDK API instance, as well as the CAN subsystem instance, are both initialised. Takes three arguments. The first argument is an integer specifying the data packet ID. The second argument is the actual data packet being transferred. The third argument is an integer specifying the data packet size. Returns 0 on success and 1 otherwise.
BMBSocialAPI.send_canbus_message(id, data, dataSize) -> int
Method called to set the target video FPS. This method must be called to pass the correct FPS value when BSocial is used to process a video, otherwise the SDK is likely to generate suboptimal predictions. There is no need to call it when BSocial is processing live camera images. This method must be called once the library API instance is initialised. Takes a single float argument fps.
BMBSocialAPI.set_video_fps(fps) -> None
Method called to pass input audio data into the instance of the SDK API. This method must be called once the SDK API instance is initialised. Takes four arguments. The first argument audio_buffer is a numpy byte array representing the input audio data. It has to be a 1-dimensional array of type np.uint8 and of shape (audio_buffer_length,), where audio_buffer_length is the number of elements stored in the array. The length of the input audio data buffer may vary. BSocial will not modify this array, leaving it the responsibility of the caller to allocate, populate and release the audio buffer passed into this method. The audio buffer may be disposed of / overwritten as soon as this method returns. The second argument audio_sample_rate is the sample rate of the input audio. The third argument audio_channels is the number of channels in the input audio. Finally, the last argument audio_format_length is the length of the audio input format in bytes. This method returns 0 on success, 1 on failure and 2 to indicate that it is still accumulating audio data required to make the first prediction.
Please note that BSocial at the moment only supports 2-byte long signed 16-bit integer audio format, single channel mono audio and a sample rate of 44100 Hz.
BMBSocialAPI.set_audio(
audio_buffer,
audio_sample_rate,
audio_channels,
audio_format_length
) -> int
Method called to pass an input image into the instance of the SDK API. This method must be called once the SDK API instance is initialised. Takes three arguments. The first argument im is a numpy array representing the input image. It has to be of type np.uint8 and of shape (image_height, image_width, image_nchannels) for colour images or alternatively (image_height, image_width) for monochrome images. The second argument im_type is a variable of type BSocialImageType, which has to be set to match the input image type. BSocial currently comes with support for the following image types:
BSocialImageType.BGRBSocialImageType.RGBBSocialImageType.BGRABSocialImageType.RGBABSocialImageType.MONOCHROMEThe third argument flip is a boolean, which can be used to specify if the image needs to be flipped vertically by the SDK API instance, which is sometimes needed for example if the image was captured directly from a camera.
BMBSocialAPI.set_image(im, im_type, flip) -> None
Method called to run inference on the input image. This method takes the image set with the set_image method, as well as audio data set with the set_audio method, and calculates output predictions for it. Must be called after the input image is passed to the SDK API instance. Returns 0 on success, 1 on inference failure, 2 on license key verification failure and 3 if the SDK was not able to locate a face on the input image.
BMBSocialAPI.run() -> int
Method called to run expressed mood estimation inference and retrieve mood estimation results. Takes no arguments. Returns a pair of values score and valid. The first value score contains the mood estimation score, which is a floating point number ranging from -1.0 to 1.0. The second value valid is a boolean flag indicating whether the returned expressed mood score is valid or not.
BMBSocialAPI.retrieve_mood() -> score, valid
Optional SDK API method which can be used to set a custom neutral mood level to alter expressed mood estimation. The neutral mood level is defined in terms of a pair of valence and arousal values, both of which are floating point numbers ranging from -1.0 to 1.0. By default, both of these values are set to zero.
BMBSocialAPI.set_neutral_mood(valence, arousal) -> None
Method which allows to change the default directory where the SDK stores neutral data of enrolled faces. Takes a single string argument path.
BMBSocialAPI.set_neutral_data_directory(path) -> None
This optional method sets a custom neutral reference pose for Action Unit (AU) calibration. By providing an image of a subject's neutral facial expression, all subsequent AU predictions are calculated relative to this specific baseline. This can significantly improve expression analysis accuracy, especially for individuals whose resting face differs from a generic model. If this function is not called, the SDK uses its default filtering and a generalized baseline for all AU calculations.
For best results, the provided image should be well-lit, front-facing, and feature the subject looking directly at the camera with a neutral expression (e.g. mouth closed, eyes open and relaxed).
It is recommended that add_neutral_face is called after init and before run. If calling add_neutral_face after data processing has started, make sure add_neutral_face and run are mutually exclusive.
Each time this is called a new persistent identity is created. These persistent identities will be saved to storage and loaded for future runs. A different neutral baseline will be applied to each identity. This method allows to provide a user-defined name.
Returns a tuple where:
BSocialAPI.add_neutral_face(
im, im_type, flip, user_defined_name
) -> (int, string)
Method which allows to add a neutral face to an existing identity using the persistent_face_id that was output at the time of creation. Returns 0 on success, 1 on error.
BSocialAPI.add_neutral_face_to_existing_identity(
im, im_type, flip, persistent_face_id_str, user_defined_name
) -> int
Remove a previously added neutral face identity using the persistent_face_id that was output at the time of creation. Returns 0 on success, 1 on error.
BSocialAPI.remove_neutral_face_with_persistent_face_id(
persistent_face_id_str
) -> int
Remove a previously added neutral face identity using the user_defined_name that was input at the time of creation. Note that user-defined names are not required to be unique. If you create multiple identities with the same user-defined name, it is undefined whether one or all of these are removed. Returns 0 on success, 1 on error.
BSocialAPI.remove_neutral_face_with_user_defined_name(
user_defined_name
) -> int
Get a list of persistent_face_id and user_defined_name pairs (in that order).
BSocialAPI.get_neutral_faces() -> list[tuple[str, str]]
Remove all added neutral face identities. Returns 0 on success, 1 on error.
BSocialAPI.clear_neutral_faces() -> int
Method called to disable post-processing of certain predictions, such as Action Unit predictions, for which time-based filtering algorithms are applied. Takes a single boolean argument enabled, which allows to control whether such post-processing takes place or not. By default the post-processing is enabled, however in rare occasions it might be beneficial to switch it off. This could for example occur while processing very short video clips, which are only several seconds long.
BMBSocialAPI.set_predictions_postprocessing_enabled(enabled) -> None
Method called to get inference results for the input image passed into the SDK API upon successful inference. Takes no arguments. Returns a list of objects of the BSocialPredictions structure type. An overview of the BSocialPredictions structure is provided in the API Outputs Reference section of this user guide. The list contains an entry per each face detected by the SDK.
BMBSocialAPI.get_predictions() -> retval
Optional SDK API method which can be used to specify the path to the output .csv file generated by the API instance containing resulting predictions. Takes a single string argument output_path.
BMBSocialAPI.set_output_path(output_path) -> None
Optional SDK API method which can be used to tell the SDK API instance when to start writing predictions to the output path specified with the aforementioned set_output_path method. This method must be called after the output path is specified. It optionally accepts a BMCSVConfig object to control which data sections are written to the CSV file. If no configuration is provided, the SDK uses the default core outputs. Returns 0 on success and 1 otherwise.
BMBSocialAPI.start_write(config: BMCSVConfig = None) -> int
Optional SDK API method which can be used to tell the SDK API instance when to stop writing predictions to the output path specified with the aforementioned set_output_path method. Takes no arguments. This method can be used to finalise predictions saving for a specific video before switching to another.
BMBSocialAPI.stop_write() -> None
Optional SDK API method which can be used to set the minimum acceptable face bounding box diagonal length in pixels. This method must be called once the SDK API instance is initialised. Takes a single positive integer argument.
BMBSocialAPI.set_min_face_diagonal(value) -> None
Optional SDK API method which can be used to load a custom emotion zones configuration file in order to alter the set of emotion zones classified by BSocial. This method must be called once the SDK API instance is initialised. Takes a single string argument emotionZonesConfigPath containing the path to the target configuration file. Returns 0 on success and 1 otherwise.
BMBSocialAPI.load_emotion_zones_config(emotionZonesConfigPath) -> int
Optional SDK API method which can be used to get the list of emotion zone labels classified by BSocial. This method must be called once the SDK API instance is initialised. Takes no arguments. Returns a list of string values.
BMBSocialAPI.get_emotion_zones() -> list[str]
Optional SDK API method which can be used to add an additional emotion zone to the set of emotion zones classified by BSocial. This method must be called once the SDK API instance is initialised. Takes five arguments. The first argument label is a string variable containing the emotion zone label to be added. The rest of the arguments are floating point values specifying the new emotion zone mean (mx and my) and standard deviation (sx and sy) parameters. Returns 0 on success and 1 otherwise.
BMBSocialAPI.add_emotion_zone(label, mx, my, sx, sy) -> int
Optional SDK API method which can be used to remove an emotion zone from the set of emotion zones classified by BSocial. This method must be called once the SDK API instance is initialised. Takes a single string argument containing the emotion zone label to be removed. Returns 0 on success and 1 otherwise.
BMBSocialAPI.remove_emotion_zone(label) -> int
Optional SDK API method primarily used for visualising predictions by overlaying an image passed to this method with inference results upon successful inference. Takes three arguments. The first argument im is a numpy array representing the overlaid image. It has to be of type np.uint8 and of shape (image_height, image_width, image_nchannels) for colour images or alternatively (image_height, image_width) for monochrome images. The second argument im_type is a variable of type BSocialImageType, which has to be set to match the input image type. The third argument flip is a boolean, which can be used to specify if the image needs to be flipped vertically by the SDK API instance. Returns a numpy array of the same shape and type as im containing the overlaid image.
BMBSocialAPI.overlay(im, im_type, flip) -> retval
The C++ API mirrors the Python API described above, with the following language-specific signatures.
std::string BSocialAPI::version();
int BSocialAPI::load_licence_key(const std::string &licenceKeyPath);
void BSocialAPI::set_licence_key(const std::string &licenceKey);
void BSocialAPI::set_log_level(const int32_t& log_level);
int32_t BSocialAPI::set_max_faces_to_track(const int32_t &max_faces_to_track);
The set of models currently included in the incremental inference mode include image quality, cough and sneeze, as well as visual voice detection.
void BSocialAPI::set_inference_increment_enabled(const bool &enabled);
void BSocialAPI::set_min_process_time(const float &time);
int BSocialAPI::init();
void BSocialAPI::reset();
int BSocialAPI::load_object_mappings(
const std::string &AOI_positions_path);
int32_t BSocialAPI::load_camera_calibration(
const std::string &camera_calibration_path);
void BSocialAPI::set_look_ahead_angles(float pitch, float yaw);
void BSocialAPI::set_camera_transform(
float x, float y, float z, float roll, float pitch, float yaw);
This method must be called once the SDK API instance is initialised. Takes a string specifying the target CAN device name, and an integer specifying the start index of the messages the CAN subsystem generates, which defaults to 0x100. Returns 0 on success and 1 otherwise.
int BSocialAPI::init_canbus(const std::string &deviceID,
int startIndex = 0x100);
int BSocialAPI::send_canbus_message(uint32_t id,
const char* data,
int32_t dataSize);
void BSocialAPI::set_video_fps(const float &fps);
int BSocialAPI::set_audio(
const unsigned char *audio_buffer,
const int& audio_buffer_length,
const int& audio_sample_rate,
const int& audio_channels,
const int& audio_format_length
);
void BSocialAPI::set_image(
unsigned short &width,
unsigned short &height,
unsigned char *data,
BSocialImageType type,
bool flip = false
);
int BSocialAPI::run();
void BSocialAPI::retrieve_mood(float& score, bool& valid);
void BSocialAPI::set_neutral_mood(float valence, float arousal);
void BSocialAPI::set_neutral_data_directory(const std::string& path);
Returns an integer status code indicating the outcome: 0 is success, 1 is failure due to insufficient image quality, 3 is failure due to no face detected.
int BSocialAPI::add_neutral_face(
unsigned short &width,
unsigned short &height,
unsigned char *data,
BSocialImageType type,
bool flip = false,
const std::string& user_defined_name,
std::string& persistant_id
);
int32_t BSocialAPI::add_neutral_face_to_existing_identity(
const std::string& persistent_face_id,
const uint16_t& width,
const uint16_t& height,
unsigned char* data,
BSocialImageType type,
bool flip,
const std::string& user_defined_name = ""
);
int BSocialAPI::remove_neutral_face_with_persistent_face_id(
const std::string& persistent_face_id
);
int BSocialAPI::remove_neutral_face_with_user_defined_name(
const std::string& user_defined_name
);
std::list<std::pair<std::string, std::string>> BSocialAPI::get_neutral_faces();
int BSocialAPI::clear_neutral_faces();
void BSocialAPI::set_predictions_postprocessing_enabled(
const bool& enabled);
void BSocialAPI::get_predictions(
std::vector<BSocialPredictions> &predictions);
void BSocialAPI::set_output_path(const std::string &output_path);
int BSocialAPI::start_write(const BMCSVConfig& config = BMCSVConfig());
void BSocialAPI::stop_write();
void BSocialAPI::set_min_face_diagonal(const unsigned short &value);
int BSocialAPI::load_emotion_zones_config(
const std::string& emotionZonesConfigPath);
void BSocialAPI::get_emotion_zones(std::vector<std::string>& emotions);
int BSocialAPI::add_emotion_zone(
const std::string& label, float mx, float my, float sx, float sy);
int BSocialAPI::remove_emotion_zone(const std::string& label);
void BSocialAPI::overlay(
unsigned short &width,
unsigned short &height,
unsigned char *data,
BSocialImageType type,
bool flip = false
);
BMCSVConfig structure allows to toggle specific data sections in the CSV output. By default, only core tracking and emotional data are enabled to reduce file size. BMCSVConfig contains the following boolean toggles which can be set to define the set of columns the SDK writes into output CSV files:
includeFrameincludeFaceIDincludeFaceDetectionincludeHeadPositionincludeHeadOrientationincludeFaceLandmarksincludeGazeincludeActionUnitsincludeAffectincludeImHOincludeImageQualityincludeDrowsiness (Automotive)includeEmotionZonesincludeVVADincludeAVAD (Core)includeSneezeCough (Automotive)includeAttentionincludeConfusion (Core)includeInteractionEngagement (Core)BSocial API returns inference results obtained from every input image passed to the SDK API instance with a variable of the dedicated BSocialPredictions structure type. This structure is composed of a number of fields, one per each of the internal BSocial components populating its respective field of the structure. An overview of each of these fields is provided further below in this section.
The predictions structure overview is presented from the C++ standpoint, however it is equally valid for the Python interface of BSocial. The Python interface outputs predictions arranged in the same way and of the same type with the exception of the fields containing C++ vector type variables, which are returned as list types in Python.
struct BSocialPredictions {
BSocialFaceID faceID;
BSocialFaceDetection faceDetection;
BSocialFaceVisibility visibility;
BSocialLandmarks landmarks;
BSocialHeadAngles headAngles;
BSocialHeadPosition headPosition;
BSocialGaze gaze;
BSocialActionUnits actionUnits;
BSocialAffect affect;
BSocialImageQuality imageQuality;
BSocialDrowsiness drowsiness; // Automotive
BSocialEmotionZones emotionZones;
BSocialVVAD vvad;
BSocialAVAD avad; // Core
BSocialSneeze sneeze; // Automotive
BSocialAttention attention;
BSocialConfusion confusion; // Core
BSocialInteractionEngagement interactionEngagement; // Core
};
BSocialFaceID is a structure type designed to hold face re-identification results composed of fields: current, count, user_defined_name, persistent_face_id and valid. The field current is an integer indicating a numerical ID of the person currently being tracked (positive integers are valid, zero is invalid). The field count is an integer indicating the total number of faces BSocial was able to identify in the current video sequence. If you have enrolled faces using the add_neutral_face function, and the face being tracked matches an enrolled face, the fields user_defined_name and persistent_face_id will be set (otherwise they are set to the empty string). The last field valid is a boolean field specifying whether face re-identification data presented in a particular instance of BSocialFaceID is valid or not.
struct BSocialFaceID {
int32_t current = 0;
int32_t count = 0;
std::string user_defined_name = "";
std::string persistent_face_id = "";
bool valid = false;
};
BSocialFaceDetection is a structure type designed to hold face detection results composed of five fields: x, y, width, height and valid. The first four fields are integers specifying face detection bounding box coordinates and dimensions. x and y contain coordinates of the top left corner of the facial bounding box. width and height contain dimensions of the facial bounding box. The last field valid is a boolean field specifying whether face detection data presented in a particular instance of BSocialFaceDetection is valid or not.
struct BSocialFaceDetection {
int32_t x = 0;
int32_t y = 0;
int32_t width = 0;
int32_t height = 0;
bool valid = false;
};
BSocialLandmarks is a structure type designed to hold facial landmarks tracking results composed of four fields: x, y, v and valid. The first two fields are vectors of floats specifying pairs of XY coordinates of every facial landmark tracked. The third field is a vector of booleans indicating visibility of every facial landmark, which could depend on either head position and / or orientation, as well as presence of occlusions. BSocial supports and outputs coordinates, as well as visibility, of 68 different facial landmarks. The last field valid is a boolean field specifying whether facial landmarks tracking data presented in a particular instance of BSocialLandmarks is valid or not.
struct BSocialLandmarks {
std::vector<float> x = std::vector<float>();
std::vector<float> y = std::vector<float>();
std::vector<bool> v = std::vector<bool>();
bool valid = false;
};
BSocialHeadAngles is a structure type designed to hold head angle estimation results composed of four fields: yaw, pitch, roll and valid. The first three fields are floats specifying head orientation angles in degrees. The last field valid is a boolean field specifying whether the head estimation angles data presented in a particular instance of BSocialHeadAngles is valid or not.
struct BSocialHeadAngles {
float yaw = 0;
float pitch = 0;
float roll = 0;
bool valid = false;
};
BSocialHeadPosition is a structure type designed to hold head position estimation results composed of four fields: x, y, z and valid. The first three fields are floats specifying absolute head position. The last field valid is a boolean field specifying whether head position data presented in a particular instance of BSocialHeadPosition is valid or not.
struct BSocialHeadPosition {
float x = 0;
float y = 0;
float z = 0;
bool valid = false;
};
BSocialFaceVisibility is a structure type designed to hold face regions visibility estimation results composed of five fields: mouth, nose, eye_left, eye_right and valid. The first four fields are booleans indicating visibility of the corresponding face regions. The last field valid is a boolean field specifying whether face regions visibility data presented in a particular instance of BSocialFaceVisibility is valid or not.
struct BSocialFaceVisibility {
bool mouth = false;
bool nose = false;
bool eye_left = false;
bool eye_right = false;
bool valid = false;
};
BSocialGaze is a structure type designed to hold social gaze direction estimation results composed of the fields angles, anglesAverage, vector, vectorAverage, leftEyeOccluded, rightEyeOccluded, leftEyeTracked, rightEyeTracked, headPoseFallback and valid. The first two fields are instances of the BSocialGazeAngles structure containing estimated social gaze angles: angles contains estimation results for the current frame, whereas anglesAverage contains results smoothed by a running mean filter. The next two fields are instances of the BSocialGazeVector structure containing estimated social gaze vector coordinates, likewise for the current frame and smoothed respectively. leftEyeOccluded and rightEyeOccluded indicate if an eye is closed or occluded by an object — if both eyes are closed then head rotation is supplied in the gaze angles. leftEyeTracked, rightEyeTracked and headPoseFallback indicate which eye(s) were tracked in a given frame and if the SDK had to fall back to head pose while attempting to resolve gaze direction. The last field valid is a boolean field specifying whether social gaze estimation data presented in a particular instance of BSocialGaze is valid or not.
struct BSocialGazeAngles {
float yaw = 0;
float pitch = 0;
};
struct BSocialGazeVector {
float x = 0;
float y = 0;
float z = 0;
};
struct BSocialGaze {
BSocialGazeAngles angles;
BSocialGazeAngles anglesAverage;
BSocialGazeVector vector;
BSocialGazeVector vectorAverage;
bool leftEyeOccluded;
bool rightEyeOccluded;
bool leftEyeTracked = false;
bool rightEyeTracked = false;
bool headPoseFallback = false;
bool valid = false;
};
BSocialActionUnits is a structure type designed to hold Action Unit intensity estimation results composed of four fields: classes, predictions, confidence and valid. The first field is a vector of strings, which contains labels of the Action Units BSocial returns intensity predictions for. The second field is a vector of floats which contains actual AU intensity predictions in the range from 0 (no activation) to 5 (maximum activation), one per each of the class labels. The third field confidence is a float defining confidence of the Action Unit predictions ranging from 0 to 1. valid is a boolean field specifying whether Action Unit intensity estimation data presented in a particular instance of BSocialActionUnits is valid or not.
struct BSocialActionUnits {
std::vector<std::string> classes = std::vector<std::string>();
std::vector<float> predictions = std::vector<float>();
float confidence = 0;
bool valid = false;
};
BSocialAffect is a structure type designed to hold affect state prediction values composed of seven fields: valence, arousal, dominance, confidence_valence, confidence_arousal, confidence_dominance and valid. The first three fields are floats containing estimated levels of Valence, Arousal and Dominance. The next three fields are floats containing confidence values of the Valence, Arousal and Dominance predictions in the range between 0 and 1. The last field valid is a boolean field specifying whether affect state levels estimation data presented in a particular instance of BSocialAffect is valid or not.
struct BSocialAffect {
float valence = 0;
float arousal = 0;
float dominance = 0;
float confidence_valence = 0;
float confidence_arousal = 0;
float confidence_dominance = 0;
bool valid = false;
};
BSocialImageQuality is a structure type designed to hold various image quality characteristics as identified by BSocial for every input image passed into the API. The structure is composed of the fields imageAcceptable, imageBlurry, imageTooDark, imageTooBright, imageIsNIR and valid. imageBlurry indicates whether the input image is not sharp enough. imageAcceptable indicates whether the input image quality was found acceptable. imageTooDark and imageTooBright indicate whether the input image is either too dark or too bright. imageIsNIR denotes if the frame is near infra-red. Finally, valid indicates whether BSocial was able to perform input image quality assessment successfully or not.
struct BSocialImageQuality {
bool imageAcceptable = false;
bool imageBlurry = false;
bool imageTooDark = false;
bool imageTooBright = false;
bool imageIsNIR = false;
bool valid = false;
};
BSocialDrowsiness (Automotive) is a structure type designed to hold drowsiness estimation results. The structure is composed of three fields: state, confidence and valid. state indicates whether the subject observed appears to be drowsy at the moment. confidence is a float indicating drowsiness estimation confidence ranging from 0 to 1. valid specifies whether drowsiness detection data presented in a particular instance of BSocialDrowsiness is valid or not.
struct BSocialDrowsiness {
bool state = false;
float confidence = 0;
bool valid = false;
};
BSocialEmotionZones is a structure type designed to hold emotion zones classification results. The structure is composed of five fields: emotions, probabilities, emotion, mood_probability_confidence and valid. emotions is an array of string values containing a list of emotion zone labels classified by the SDK. probabilities is an array of floating point values ranging from 0 to 1 containing classification probabilities of each emotion zone listed in the emotions array. emotion is a string variable containing the emotion zone label classified as the closest representation of the subject's present emotional state, that is the emotion zone with the highest classification probability score. mood_probability_confidence holds prediction confidence of emotion zones estimation. The last field valid is a boolean specifying whether emotion zones classification data presented in a particular instance of BSocialEmotionZones is valid or not.
struct BSocialEmotionZones {
std::vector<std::string> emotions = std::vector<std::string>();
std::vector<float> probabilities = std::vector<float>();
std::string emotion = "Unknown";
float mood_probability_confidence = 0;
bool valid = false;
};
BSocialVVAD is a structure type designed to hold visual voice detection results. The structure is composed of three fields: state, confidence and valid. state indicates whether the subject observed appears to be talking at the moment or not. confidence is a floating point value between 0 and 1 where 0 denotes that the model is predicting that the subject is not talking and 1 that they are talking. valid specifies whether visual voice detection data presented in a particular instance of BSocialVVAD is valid or not.
struct BSocialVVAD {
bool state = false;
float confidence = 0.0f;
bool valid = false;
};
BSocialAVAD (Core) is a structure type designed to hold audio voice detection results, with the same field layout as BSocialVVAD.
struct BSocialAVAD {
bool state = false;
float confidence = 0.0f;
bool valid = false;
};
BSocialSneeze (Automotive) is a structure type designed to hold sneeze and cough detection results. The structure is composed of six fields: cough, sneeze, confidence_negative, confidence_cough, confidence_sneeze and valid. The first two fields are booleans indicating cough and sneeze activation respectively. The following three fields are floats indicating confidence values of the cough and sneeze detection ranging from 0 to 1. The last field valid specifies whether sneeze and cough detection data presented in a particular instance of BSocialSneeze is valid or not.
struct BSocialSneeze {
bool cough = false;
bool sneeze = false;
float confidence_negative = 0;
float confidence_cough = 0;
float confidence_sneeze = 0;
bool valid = false;
};
BSocialAttention is a structure type designed to hold gaze-to-object mapping results. The structure is composed of looked_at_object, looking_straight_ahead, gaze_origin_x/y/z, gaze_vector_x/y/z, projected_intersection_x/y and valid. looked_at_object is a string indicating which object or area the subject observed appears to be looking at. looking_straight_ahead indicates whether the subject appears to be looking ahead. The remaining float fields contain 3D coordinates of gaze origin, gaze vector direction, as well as 2D coordinates of the point marking intersection of the gaze vector with the attention area of interest (AOI) in the AOI coordinate space. valid specifies whether gaze-to-object mapping data presented in a particular instance of BSocialAttention is valid or not.
struct BSocialAttention {
std::string looked_at_object = "";
bool looking_straight_ahead = false;
float gaze_origin_x = 0.0f;
float gaze_origin_y = 0.0f;
float gaze_origin_z = 0.0f;
float gaze_vector_x = 0.0f;
float gaze_vector_y = 0.0f;
float gaze_vector_z = 0.0f;
float projected_intersection_x = 0.0f;
float projected_intersection_y = 0.0f;
bool valid = false;
};
BSocialConfusion (Core) is a structure type that holds the confusion estimation. The structure consists of two fields: confusion and valid. confusion is the confusion estimation score that ranges from 0 to 1, where 0 is not confused and 1 is completely confused. valid indicates if the confusion estimation is valid for this frame.
struct BSocialConfusion {
float confusion = 0.0f;
bool valid = false;
};
BSocialInteractionEngagement (Core) is a structure type that holds the interaction engagement score and a flag denoting the intent to engage. The structure consists of three fields: intention_to_interact, overall_engagement_score and valid.
struct BSocialInteractionEngagement {
bool intention_to_interact = false;
float overall_engagement_score = 0.0f;
bool valid = false;
};
AI system provider: Blueskeye Ltd
Company registration number: 11953581
Contact e-mail: help@blueskeye.com
Blueskeye Ltd Headquarters
The Ingenuity Centre
University of Nottingham Innovation Park
Triumph Road
Nottingham
NG7 2TU, UK